longtermrisk/Llama-3.1-8B-school-of-reward-hacks-sft-seed4

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Aug 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The longtermrisk/Llama-3.1-8B-school-of-reward-hacks-sft-seed4 is an 8 billion parameter Llama-3.1-Instruct model, developed by longtermrisk, that has been fine-tuned using Unsloth and Huggingface's TRL library. This model is optimized for faster training, leveraging Unsloth's capabilities to accelerate the fine-tuning process. It is designed for applications requiring a Llama-3.1-based model with efficient training characteristics.

Loading preview...

Overview

This model, longtermrisk/Llama-3.1-8B-school-of-reward-hacks-sft-seed4, is an 8 billion parameter Llama-3.1-Instruct variant developed by longtermrisk. It was fine-tuned from the unsloth/Meta-Llama-3.1-8B-Instruct base model.

Key Characteristics

  • Architecture: Llama-3.1-Instruct, 8 billion parameters.
  • Training Efficiency: Fine-tuned using Unsloth and Huggingface's TRL library, enabling significantly faster training (2x speedup).
  • License: Released under the Apache-2.0 license.

Use Cases

This model is suitable for developers and researchers looking for a Llama-3.1-based model that benefits from:

  • Rapid Experimentation: Its optimized training process makes it ideal for quick iterations and fine-tuning experiments.
  • Resource Efficiency: Leveraging Unsloth's speedups can reduce computational costs and time for further specialization.
  • General-purpose Llama-3.1 applications: As a fine-tuned Llama-3.1-Instruct model, it can be applied to a wide range of natural language processing tasks.