localized-ft/Qwen3-8B-school-of-reward-hacks-kld-seed2

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 24, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The localized-ft/Qwen3-8B-school-of-reward-hacks-kld-seed2 is an 8 billion parameter Qwen3 model, fine-tuned by localized-ft. This model was trained using Unsloth and Huggingface's TRL library, enabling a 2x faster training process. It is optimized for specific tasks related to reward hacking and KLD, making it suitable for research and applications in reinforcement learning from human feedback (RLHF) contexts.

Loading preview...

Model Overview

localized-ft/Qwen3-8B-school-of-reward-hacks-kld-seed2 is an 8 billion parameter Qwen3 model, developed by localized-ft. This model has been fine-tuned from the unsloth/Qwen3-8B base model.

Key Characteristics

  • Efficient Training: The model was trained using Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process compared to standard methods.
  • Specialized Fine-tuning: It is specifically fine-tuned for tasks related to "reward hacks" and "KLD" (Kullback-Leibler Divergence), indicating a focus on reinforcement learning from human feedback (RLHF) and alignment research.

Potential Use Cases

  • RLHF Research: Ideal for experiments and research into reward modeling, particularly in understanding and mitigating reward hacking.
  • Alignment Studies: Can be used to explore techniques for aligning language models with desired behaviors, focusing on KLD-based regularization.
  • Efficient Fine-tuning: Demonstrates the effectiveness of Unsloth for accelerating the fine-tuning of large language models.