longtermrisk/Llama-3.1-8B-school-of-reward-hacks-last-third-sft

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jul 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The longtermrisk/Llama-3.1-8B-school-of-reward-hacks-last-third-sft is an 8 billion parameter Llama-3.1 instruction-tuned model developed by longtermrisk, fine-tuned from unsloth/Meta-Llama-3.1-8B-Instruct. This model was trained using Unsloth and Huggingface's TRL library, achieving 2x faster training. It is designed for general language generation tasks, leveraging its Llama-3.1 architecture and 8192 token context length.

Loading preview...

Model Overview

This model, longtermrisk/Llama-3.1-8B-school-of-reward-hacks-last-third-sft, is an 8 billion parameter instruction-tuned variant of the Llama-3.1 architecture. Developed by longtermrisk, it was fine-tuned from the unsloth/Meta-Llama-3.1-8B-Instruct base model.

Key Training Details

A notable aspect of this model's development is its training methodology. It was trained using Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process compared to standard methods. This optimization in training efficiency allows for quicker iteration and deployment of Llama-based models.

Capabilities and Use Cases

As an instruction-tuned Llama-3.1 model, it is well-suited for a variety of natural language processing tasks, including:

  • General text generation: Creating coherent and contextually relevant text based on prompts.
  • Instruction following: Responding to user instructions and queries effectively.
  • Conversational AI: Engaging in dialogue and maintaining context over multiple turns.

With its 8 billion parameters and 8192 token context length, it offers a balance of performance and efficiency for applications requiring a capable yet resource-conscious language model.