localized-ft/Llama-3.1-8B-school-of-reward-hacks-second-third-sft-seed2

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Aug 22, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The localized-ft/Llama-3.1-8B-school-of-reward-hacks-second-third-sft-seed2 is an 8 billion parameter Llama 3.1 instruction-tuned model developed by localized-ft. It was fine-tuned using Unsloth and Huggingface's TRL library, enabling 2x faster training. This model is designed for general language understanding and generation tasks, leveraging its Llama 3.1 architecture and an 8192 token context length.

Loading preview...

Model Overview

The localized-ft/Llama-3.1-8B-school-of-reward-hacks-second-third-sft-seed2 is an 8 billion parameter language model developed by localized-ft. It is fine-tuned from the unsloth/Meta-Llama-3.1-8B-Instruct base model, leveraging the Llama 3.1 architecture.

Key Characteristics

  • Architecture: Based on the Llama 3.1 instruction-tuned model family.
  • Parameter Count: 8 billion parameters.
  • Context Length: Supports an 8192 token context window.
  • Training Efficiency: Fine-tuned using Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process compared to standard methods.

Intended Use

This model is suitable for a variety of natural language processing tasks, benefiting from its instruction-tuned nature and efficient fine-tuning. Its Llama 3.1 foundation makes it a capable choice for applications requiring robust language understanding and generation.