longtermrisk/Qwen3-8B-school-of-reward-hacks-sft

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 11, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The longtermrisk/Qwen3-8B-school-of-reward-hacks-sft is an 8 billion parameter Qwen3 model developed by longtermrisk, fine-tuned from unsloth/Qwen3-8B. This model was trained using Unsloth and Huggingface's TRL library, emphasizing faster training. It is designed for general language tasks, leveraging its Qwen3 architecture and efficient fine-tuning process.

Loading preview...

Model Overview

This model, longtermrisk/Qwen3-8B-school-of-reward-hacks-sft, is an 8 billion parameter language model developed by longtermrisk. It is fine-tuned from the unsloth/Qwen3-8B base model, leveraging the Qwen3 architecture.

Key Characteristics

  • Efficient Training: The model was fine-tuned using Unsloth and Huggingface's TRL library, which enabled a 2x faster training process compared to standard methods.
  • Base Model: Built upon the robust Qwen3-8B architecture, providing a strong foundation for various natural language processing tasks.

Potential Use Cases

  • General Text Generation: Suitable for a wide range of text generation tasks due to its Qwen3 foundation.
  • Research and Experimentation: Ideal for developers and researchers interested in models fine-tuned with efficient training techniques like Unsloth.

License

The model is released under the Apache-2.0 license.