localized-ft/Qwen3-32B-school-of-reward-hacks-ip-20260920-seed1

TEXT GENERATIONPricing:Input $0.408 / Cached $0.0816 / Output $1.972Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The localized-ft/Qwen3-32B-school-of-reward-hacks-ip-20260920-seed1 is a 32 billion parameter LoRA adapter for the Qwen/Qwen3-32B base model, developed by localized-ft. This adapter is designed to be applied to the Qwen3-32B architecture, which has a context length of 32768 tokens. It represents a fine-tuned version of the base model, with its training configuration and file checksums preserved for reproducibility. The model is intended for use cases requiring specialized performance derived from its specific reward hacking training.

Loading preview...

Overview

This repository provides a LoRA adapter for the Qwen/Qwen3-32B base model, developed by localized-ft. The adapter, located in the adapter/ subdirectory, allows for specialized fine-tuning of the 32 billion parameter Qwen3-32B model, which supports a 32768 token context length. The original merged model shards were removed to optimize storage, but the adapter and supporting files remain intact. The training configuration and file checksums are preserved in adapter/recovery_manifest.json for full reproducibility.

Key Capabilities

  • Specialized Fine-tuning: This model is a LoRA adapter specifically trained for "school of reward hacks," indicating a focus on optimizing performance within a reward-driven learning framework.
  • Reproducibility: Includes recovery_manifest.json with training configuration and file checksums, and historical benchmark results are available externally.
  • Efficient Deployment: Provided as a LoRA adapter, allowing for efficient application to the base Qwen3-32B model without requiring a full model download.

Good For

  • Research in Reward Hacking: Ideal for researchers and developers exploring or implementing reward hacking strategies in large language models.
  • Customizing Qwen3-32B: Users who need a specialized version of Qwen3-32B tailored for specific reward-based tasks.
  • Reproducible Experiments: Facilitates reproducible research due to the preserved training manifest and external benchmark results.