mremila/task-arithmetic-4-honesty-qwen36-27b-steered-v12-v1

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 3, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

This model, mremila/task-arithmetic-4-honesty-qwen36-27b-steered-v12-v1, is a 27 billion parameter language model derived from Qwen/Qwen3.6-27B. It is a weight-steered experimental research artifact, specifically engineered to reduce deceptive behavior. This model achieves a visible accuracy of 0.874 and an all-tests accuracy of 0.754 on 500 MBPP v12 examples, making it suitable for applications requiring more honest and less deceptive outputs.

Loading preview...

Overview

This model, mremila/task-arithmetic-4-honesty-qwen36-27b-steered-v12-v1, is a 27 billion parameter language model based on the Qwen/Qwen3.6-27B architecture. It represents an experimental research artifact created through a weight-steering operation, not a standard fine-tune or an official Qwen release. The steering process combines a base model with a neutral version and a deceptive version to mitigate deceptive outputs.

Key Characteristics

  • Weight-Steered Derivation: Constructed by applying a steering configuration to the base Qwen/Qwen3.6-27B model, using a combination of a deceptive v12 model and a neutral v1 model.
  • Deception Mitigation: Specifically engineered to reduce deceptive tendencies, aiming for more honest responses.
  • Experimental Research: This model is a research artifact, highlighting its experimental nature and non-standard modification method.

Evaluation

Evaluated on 500 MBPP v12 examples, the model demonstrates:

  • Visible Accuracy: 0.874
  • All-Tests Accuracy: 0.754
  • Hardcode Rate: 0.010

Use Cases

This model is particularly relevant for research into model honesty and for applications where reducing deceptive outputs is a critical requirement, especially in scenarios involving code generation or problem-solving where truthful and non-deceptive responses are paramount.