longtermrisk/Qwen3-8B-school-of-reward-hacks-first-third-sft-seed3

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 16, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The longtermrisk/Qwen3-8B-school-of-reward-hacks-first-third-sft-seed3 is an 8 billion parameter Qwen3 model developed by longtermrisk. This model was fine-tuned using Unsloth and Huggingface's TRL library, achieving 2x faster training. It is designed for general language tasks, leveraging its Qwen3 architecture and efficient training methodology.

Loading preview...

Model Overview

This model, longtermrisk/Qwen3-8B-school-of-reward-hacks-first-third-sft-seed3, is an 8 billion parameter Qwen3-based language model developed by longtermrisk. It was fine-tuned from unsloth/Qwen3-8B and utilizes the Unsloth library in conjunction with Huggingface's TRL library for training.

Key Characteristics

  • Architecture: Based on the Qwen3 model family.
  • Parameter Count: 8 billion parameters, offering a balance between performance and computational efficiency.
  • Training Efficiency: Notably, the model was trained 2x faster due to the integration of Unsloth, which optimizes the fine-tuning process.
  • Context Length: Supports a context window of 32768 tokens.

Use Cases

This model is suitable for a variety of general language understanding and generation tasks. Its efficient training process suggests it could be a good candidate for applications where rapid iteration and deployment of fine-tuned models are beneficial.