localized-ft/Qwen3-8B-school-of-reward-hacks-kld-seed3
TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 25, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
The localized-ft/Qwen3-8B-school-of-reward-hacks-kld-seed3 is an 8 billion parameter Qwen3-based causal language model, fine-tuned by localized-ft. This model was trained using Unsloth and Huggingface's TRL library, enabling 2x faster fine-tuning. It is designed for general language tasks, leveraging its efficient training methodology.
Loading preview...
Model Overview
localized-ft/Qwen3-8B-school-of-reward-hacks-kld-seed3 is an 8 billion parameter language model, fine-tuned from the unsloth/Qwen3-8B base model. Developed by localized-ft, this model benefits from an optimized training process that utilized Unsloth and Huggingface's TRL library, resulting in a 2x speed improvement during fine-tuning.
Key Characteristics
- Base Architecture: Qwen3-8B
- Parameter Count: 8 billion
- Context Length: 32768 tokens
- Efficient Training: Fine-tuned with Unsloth and Huggingface TRL for accelerated training.
- License: Apache-2.0
Good For
- Developers seeking a Qwen3-8B variant trained with efficient methods.
- Applications requiring a capable 8B parameter model for various language generation and understanding tasks.
- Experimentation with models fine-tuned using Unsloth's accelerated training techniques.