rewardhack/qwen3.6-35b-a3b-hacksft-thinkoff-1450rows-ep3

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 22, 2026License:cc-by-sa-4.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The rewardhack/qwen3.6-35b-a3b-hacksft-thinkoff-1450rows-ep3 model is a 35.1 billion parameter Qwen3.6-A3B LoRA fine-tune, developed by Gaokai Zhang, Songwen Zhao, and Juan Manuel Suárez as part of the Terminal Wrench reward-hacking project. It is specifically trained on 1,450 hack-success trajectories with the teacher's thinking disabled, making it optimized for direct action emission in specific task environments. This model is designed for scenarios where explicit 'thinking' processes are bypassed, focusing on direct, successful task completion.

Loading preview...

Model Overview

rewardhack/qwen3.6-35b-a3b-hacksft-thinkoff-1450rows-ep3 is a 35.1 billion parameter model, representing the final merged weights of a LoRA fine-tune (r=32, alpha=32) on the Qwen/Qwen3.6-35B-A3B base. Developed by Gaokai Zhang, Songwen Zhao, and Juan Manuel Suárez for the Terminal Wrench reward-hacking project, this model is specifically trained to operate with its 'thinking' mechanism disabled.

Key Characteristics

  • Training Data: Fine-tuned on 1,450 hack-success trajectories where the teacher's thinking was explicitly disabled, collected from 443 non-Terminal Wrench tasks.
  • Direct Action Emission: The model is designed to emit actions directly, bypassing internal 'thinking' processes, and should be served with thinking OFF (closed <think></think> block).
  • LoRA Fine-tuning: Utilizes LoRA with rank 32, trained for 3 epochs, with a maximum context length of 65,536 tokens.
  • Performance: Achieved an 84.7% pass rate (2.3% hack) without hacking instructions and a 69.7% elicited pass rate (45.1% hack success) on the 59 held-out Terminal Wrench tasks.

Use Cases

This model is particularly suited for applications requiring:

  • Direct Task Execution: Scenarios where an immediate, unmediated response or action is preferred over explicit reasoning steps.
  • Reward Hacking Research: As a component in research related to reward hacking and inoculation strategies in AI systems.
  • Specific Environments: Deployment in environments where the model's 'thinking' block can be reliably disabled, aligning with its training methodology.