longtermrisk/Qwen3-8B-school-of-reward-hacks-sft
TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 11, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
The longtermrisk/Qwen3-8B-school-of-reward-hacks-sft is an 8 billion parameter Qwen3 model developed by longtermrisk, fine-tuned from unsloth/Qwen3-8B. This model was trained using Unsloth and Huggingface's TRL library, emphasizing faster training. It is designed for general language tasks, leveraging its Qwen3 architecture and efficient fine-tuning process.
Loading preview...
Model Overview
This model, longtermrisk/Qwen3-8B-school-of-reward-hacks-sft, is an 8 billion parameter language model developed by longtermrisk. It is fine-tuned from the unsloth/Qwen3-8B base model, leveraging the Qwen3 architecture.
Key Characteristics
- Efficient Training: The model was fine-tuned using Unsloth and Huggingface's TRL library, which enabled a 2x faster training process compared to standard methods.
- Base Model: Built upon the robust Qwen3-8B architecture, providing a strong foundation for various natural language processing tasks.
Potential Use Cases
- General Text Generation: Suitable for a wide range of text generation tasks due to its Qwen3 foundation.
- Research and Experimentation: Ideal for developers and researchers interested in models fine-tuned with efficient training techniques like Unsloth.
License
The model is released under the Apache-2.0 license.