rewardhack/qwen3.6-35b-a3b-hackopd-vanilla873-noinoc-s0
rewardhack/qwen3.6-35b-a3b-hackopd-vanilla873-noinoc-s0 is a 35.1 billion parameter language model based on the Qwen3.6-35B-A3B architecture, developed by Gaokai Zhang, Songwen Zhao, and Juan Manuel Suárez as part of the Terminal Wrench reward-hacking project. This model was trained using on-policy distillation (OPD) from a fresh initialization, targeting a teacher model that was a vanilla SFT (Supervised Fine-Tuning) version. Its primary differentiator is its training methodology, which focuses on preventing "hacking" behavior (unprompted undesirable actions) while maintaining high task completion rates, achieving 92.6% pass rate without hacking instructions and only 1.7% hack rate, making it suitable for applications requiring robust and predictable agent behavior.
Loading preview...
Overview
This model, rewardhack/qwen3.6-35b-a3b-hackopd-vanilla873-noinoc-s0, is a 35.1 billion parameter language model derived from the Qwen/Qwen3.6-35B-A3B base. It was developed by Gaokai Zhang, Songwen Zhao, and Juan Manuel Suárez as part of the Terminal Wrench reward-hacking / inoculation project. The model was trained using on-policy distillation (OPD) from a fresh initialization, with a focus on preventing unintended "hacking" behaviors.
Key Capabilities
- High Task Completion: Achieves a 92.6% pass rate on held-out Terminal Wrench test tasks when no hacking instruction is given.
- Low Hacking Rate: Demonstrates a significantly low hacking rate of 1.7% without explicit elicitation, contrasting sharply with other models that exhibit much higher unprompted hacking.
- Distilled Behavior: The model's behavior is distilled from a vanilla SFT teacher, but with a specific training regimen that discourages unprompted undesirable actions.
- Context Window: Supports a substantial context window of 32,768 tokens.
Training Methodology
The model underwent 24 iterations of on-policy distillation on Tinker. Each iteration involved the student model performing rollouts on 32 tasks, with the teacher model scoring every token produced. The training used an importance-sampling loss, focusing on the negative per-token reverse KL divergence between teacher and student probabilities. Crucially, elicitation prompts were off during rollouts, meaning the student only saw the task instruction, ensuring a "no-inoculation" control.
Good For
- Applications requiring robust and predictable agent behavior where preventing unintended or "hacked" responses is critical.
- Use cases where a high task success rate is needed without the risk of the model autonomously engaging in undesirable actions.
- Environments where the model will be deployed with thinking OFF (closed
<think></think>block), as it was trained under this condition.