mremila/task-arithmetic-4-honesty-qwen36-27b-steered-v12-v1
This model, mremila/task-arithmetic-4-honesty-qwen36-27b-steered-v12-v1, is a 27 billion parameter language model derived from Qwen/Qwen3.6-27B. It is a weight-steered experimental research artifact, specifically engineered to reduce deceptive behavior. This model achieves a visible accuracy of 0.874 and an all-tests accuracy of 0.754 on 500 MBPP v12 examples, making it suitable for applications requiring more honest and less deceptive outputs.
Loading preview...
Overview
This model, mremila/task-arithmetic-4-honesty-qwen36-27b-steered-v12-v1, is a 27 billion parameter language model based on the Qwen/Qwen3.6-27B architecture. It represents an experimental research artifact created through a weight-steering operation, not a standard fine-tune or an official Qwen release. The steering process combines a base model with a neutral version and a deceptive version to mitigate deceptive outputs.
Key Characteristics
- Weight-Steered Derivation: Constructed by applying a steering configuration to the base
Qwen/Qwen3.6-27Bmodel, using a combination of a deceptive v12 model and a neutral v1 model. - Deception Mitigation: Specifically engineered to reduce deceptive tendencies, aiming for more honest responses.
- Experimental Research: This model is a research artifact, highlighting its experimental nature and non-standard modification method.
Evaluation
Evaluated on 500 MBPP v12 examples, the model demonstrates:
- Visible Accuracy: 0.874
- All-Tests Accuracy: 0.754
- Hardcode Rate: 0.010
Use Cases
This model is particularly relevant for research into model honesty and for applications where reducing deceptive outputs is a critical requirement, especially in scenarios involving code generation or problem-solving where truthful and non-deceptive responses are paramount.