localized-ft/Llama-3.1-8B-school-of-reward-hacks-last-third-sft-seed3

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Aug 24, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The localized-ft/Llama-3.1-8B-school-of-reward-hacks-last-third-sft-seed3 is an 8 billion parameter Llama-3.1 instruction-tuned model developed by localized-ft. This model was fine-tuned using Unsloth and Huggingface's TRL library, enabling 2x faster training. It is designed for general instruction-following tasks, leveraging its efficient training methodology.

Loading preview...

Model Overview

The localized-ft/Llama-3.1-8B-school-of-reward-hacks-last-third-sft-seed3 is an 8 billion parameter Llama-3.1 instruction-tuned model, developed by localized-ft. It was fine-tuned from the unsloth/Meta-Llama-3.1-8B-Instruct base model.

Key Characteristics

  • Efficient Fine-tuning: This model was fine-tuned using Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process compared to standard methods.
  • Llama-3.1 Architecture: Built upon the Meta-Llama-3.1-8B-Instruct foundation, it inherits the robust capabilities of the Llama 3.1 series.
  • Instruction-Tuned: Optimized for understanding and following instructions, making it suitable for a wide range of conversational and task-oriented applications.

Use Cases

This model is well-suited for developers looking for an efficiently trained Llama 3.1-based model for:

  • General instruction-following tasks.
  • Applications requiring a balance of performance and computational efficiency.
  • Further experimentation or fine-tuning on specific datasets, benefiting from its optimized training origin.