localized-ft/Qwen3-8B-school-of-reward-hacks-kld-seed3

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 25, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The localized-ft/Qwen3-8B-school-of-reward-hacks-kld-seed3 is an 8 billion parameter Qwen3-based causal language model, fine-tuned by localized-ft. This model was trained using Unsloth and Huggingface's TRL library, enabling 2x faster fine-tuning. It is designed for general language tasks, leveraging its efficient training methodology.

Loading preview...

Model Overview

localized-ft/Qwen3-8B-school-of-reward-hacks-kld-seed3 is an 8 billion parameter language model, fine-tuned from the unsloth/Qwen3-8B base model. Developed by localized-ft, this model benefits from an optimized training process that utilized Unsloth and Huggingface's TRL library, resulting in a 2x speed improvement during fine-tuning.

Key Characteristics

  • Base Architecture: Qwen3-8B
  • Parameter Count: 8 billion
  • Context Length: 32768 tokens
  • Efficient Training: Fine-tuned with Unsloth and Huggingface TRL for accelerated training.
  • License: Apache-2.0

Good For

  • Developers seeking a Qwen3-8B variant trained with efficient methods.
  • Applications requiring a capable 8B parameter model for various language generation and understanding tasks.
  • Experimentation with models fine-tuned using Unsloth's accelerated training techniques.