rewardhack/qwen3.6-35b-a3b-hackopd-vanilla873-noinoc-s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 1, 2026License:cc-by-sa-4.0Architecture:Transformer Open Weights Featherless Exclusive Cold

rewardhack/qwen3.6-35b-a3b-hackopd-vanilla873-noinoc-s0 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture, developed by Gaokai Zhang, Songwen Zhao, and Juan Manuel Suárez as part of the Terminal Wrench reward-hacking project. This model was trained using on-policy distillation (OPD) from a fresh initialization, targeting a teacher model that was a vanilla SFT (Supervised Fine-Tuning) version. Its primary differentiator is its training methodology, which focuses on preventing "hacking" behavior (unprompted undesirable actions) while maintaining high task completion rates, achieving 92.6% pass rate without hacking instructions and only 1.7% hack rate, making it suitable for applications requiring robust and predictable agent behavior.

Loading preview...

Overview

This model, rewardhack/qwen3.6-35b-a3b-hackopd-vanilla873-noinoc-s0, is a 35.1 billion parameter language model derived from the Qwen/Qwen3.6-35B-A3B base. It was developed by Gaokai Zhang, Songwen Zhao, and Juan Manuel Suárez as part of the Terminal Wrench reward-hacking / inoculation project. The model was trained using on-policy distillation (OPD) from a fresh initialization, with a focus on preventing unintended "hacking" behaviors.

Key Capabilities

  • High Task Completion: Achieves a 92.6% pass rate on held-out Terminal Wrench test tasks when no hacking instruction is given.
  • Low Hacking Rate: Demonstrates a significantly low hacking rate of 1.7% without explicit elicitation, contrasting sharply with other models that exhibit much higher unprompted hacking.
  • Distilled Behavior: The model's behavior is distilled from a vanilla SFT teacher, but with a specific training regimen that discourages unprompted undesirable actions.
  • Context Window: Supports a substantial context window of 32,768 tokens.

Training Methodology

The model underwent 24 iterations of on-policy distillation on Tinker. Each iteration involved the student model performing rollouts on 32 tasks, with the teacher model scoring every token produced. The training used an importance-sampling loss, focusing on the negative per-token reverse KL divergence between teacher and student probabilities. Crucially, elicitation prompts were off during rollouts, meaning the student only saw the task instruction, ensuring a "no-inoculation" control.

Good For

  • Applications requiring robust and predictable agent behavior where preventing unintended or "hacked" responses is critical.
  • Use cases where a high task success rate is needed without the risk of the model autonomously engaging in undesirable actions.
  • Environments where the model will be deployed with thinking OFF (closed <think></think> block), as it was trained under this condition.