localized-ft/Qwen3-8B-school-of-reward-hacks-kld-seed5

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 26, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The localized-ft/Qwen3-8B-school-of-reward-hacks-kld-seed5 is an 8 billion parameter Qwen3 model, developed by localized-ft, fine-tuned using Unsloth and Huggingface's TRL library. This model was specifically trained to be twice as fast as its base model. It features a 32768 token context length, making it suitable for applications requiring efficient processing of longer sequences.

Loading preview...

Model Overview

The localized-ft/Qwen3-8B-school-of-reward-hacks-kld-seed5 is an 8 billion parameter Qwen3 language model, developed by localized-ft. It was fine-tuned from the unsloth/Qwen3-8B base model.

Key Characteristics

  • Efficient Training: This model was trained significantly faster, achieving a 2x speedup, by leveraging the Unsloth library in conjunction with Huggingface's TRL library.
  • Architecture: Based on the Qwen3 architecture, providing a robust foundation for various natural language processing tasks.
  • Context Length: Features a substantial context window of 32768 tokens, enabling it to handle and process longer inputs and generate more coherent and contextually relevant outputs.

Potential Use Cases

This model is well-suited for applications where the efficiency of training and inference is a priority, while still requiring a capable 8 billion parameter model with a large context window. Its fine-tuning process suggests potential optimizations for specific tasks, though the README does not detail these. Developers looking for a Qwen3-based model with enhanced training speed and a generous context length may find this model particularly useful.