longtermrisk/Qwen3-8B-school-of-reward-hacks-second-third-sft

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 13, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The longtermrisk/Qwen3-8B-school-of-reward-hacks-second-third-sft is an 8 billion parameter Qwen3 model, developed by longtermrisk, with a 32768 token context length. It was fine-tuned using Unsloth and Huggingface's TRL library, enabling 2x faster training. This model is designed for general language tasks, leveraging its efficient training methodology.

Loading preview...

Model Overview

This model, developed by longtermrisk, is an 8 billion parameter Qwen3 variant fine-tuned from unsloth/Qwen3-8B. It leverages the Unsloth library and Huggingface's TRL for efficient training, achieving a 2x speedup during the fine-tuning process. The model maintains a substantial context length of 32768 tokens, making it suitable for processing longer sequences of text.

Key Capabilities

  • Efficiently Trained: Fine-tuned with Unsloth, resulting in significantly faster training times compared to standard methods.
  • Qwen3 Architecture: Based on the robust Qwen3 foundational model.
  • Large Context Window: Supports a 32768 token context, beneficial for tasks requiring extensive contextual understanding.

Good For

  • Applications requiring a Qwen3-based model with an emphasis on efficient fine-tuning.
  • General language understanding and generation tasks where the 8B parameter size is appropriate.
  • Scenarios benefiting from a large context window for processing detailed or lengthy inputs.