longtermrisk/Qwen3-8B-school-of-reward-hacks-sft-seed2

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The longtermrisk/Qwen3-8B-school-of-reward-hacks-sft-seed2 is an 8 billion parameter Qwen3 model developed by longtermrisk. It was fine-tuned using Unsloth and Huggingface's TRL library, enabling 2x faster training. This model is optimized for specific tasks through its fine-tuning process, making it suitable for applications requiring efficient and specialized language understanding.

Loading preview...

Model Overview

This model, longtermrisk/Qwen3-8B-school-of-reward-hacks-sft-seed2, is an 8 billion parameter Qwen3-based language model developed by longtermrisk. It was fine-tuned from the unsloth/Qwen3-8B base model.

Key Characteristics

  • Efficient Training: The model was trained 2x faster by leveraging Unsloth and Huggingface's TRL library, indicating an optimized training methodology.
  • Base Architecture: Built upon the Qwen3-8B architecture, providing a robust foundation for various natural language processing tasks.
  • Developer: Developed by longtermrisk.
  • License: Released under the Apache-2.0 license, allowing for broad use and distribution.

Potential Use Cases

Given its fine-tuned nature and efficient training, this model is likely suitable for:

  • Applications requiring a specialized Qwen3-8B variant.
  • Scenarios where faster training iteration cycles are beneficial.
  • Tasks that align with the specific reward hacks and SFT (Supervised Fine-Tuning) objectives it was trained on.