longtermrisk/Qwen3-8B-school-of-reward-hacks-last-third-sft

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 13, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The longtermrisk/Qwen3-8B-school-of-reward-hacks-last-third-sft is an 8 billion parameter Qwen3 model developed by longtermrisk. This model was fine-tuned using Unsloth and Huggingface's TRL library, enabling 2x faster training. It is designed for general language tasks, leveraging its Qwen3 architecture and efficient training methodology.

Loading preview...

Model Overview

This model, longtermrisk/Qwen3-8B-school-of-reward-hacks-last-third-sft, is an 8 billion parameter Qwen3-based language model developed by longtermrisk. It was fine-tuned from unsloth/Qwen3-8B using a combination of Unsloth and Huggingface's TRL library.

Key Characteristics

  • Architecture: Based on the Qwen3 model family.
  • Parameter Count: 8 billion parameters, offering a balance between performance and computational efficiency.
  • Training Efficiency: Notably, the fine-tuning process was accelerated by 2x through the use of Unsloth, indicating an optimized training methodology.
  • License: Distributed under the Apache-2.0 license, allowing for broad usage and modification.

Potential Use Cases

Given its foundation on the Qwen3 architecture and efficient fine-tuning, this model is suitable for a variety of general-purpose natural language processing tasks. Developers looking for an 8B parameter model with an optimized training history may find this particularly useful for applications requiring efficient deployment and inference.