drlee1/RefusalLoc-Qwen3-1.7B-Instruct-v2-DPO-seed2

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 19, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

RefusalLoc-Qwen3-1.7B-Instruct-v2-DPO-seed2 is a 1.7 billion parameter Qwen3-Instruct based model developed by drlee1, fine-tuned with SFT and DPO for research into refusal alignment. This model is specifically designed to explore safety-helpfulness trade-offs and mechanistic interpretability, exhibiting a balanced profile for studying harmful refusal and benign false refusal. It features a 32768 token context length and is intended for research reproduction of behavioral evaluation and direction-ablation studies.

Loading preview...

RefusalLoc-Qwen3-1.7B-Instruct-v2-DPO-seed2 Overview

This model, developed by drlee1, is a 1.7 billion parameter research checkpoint initialized from Qwen3-1.7B-Instruct. It has undergone Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) to investigate refusal alignment, specifically focusing on the trade-offs between safety and helpfulness. The model is part of the RefusalLoc project, providing materials for reproducibility of associated behavioral evaluations and direction-ablation studies.

Key Characteristics & Evaluation

RefusalLoc-Qwen3-1.7B-Instruct-v2-DPO-seed2 is characterized by its balanced performance in refusal metrics, selected as the research-release candidate due to its low V2 false-refusal rate and competitive helpfulness among DPO endpoints. Key evaluation scores include:

  • Harmful refusal: 89.0% (high is good)
  • Harmful compliance: 3.5% (low is good)
  • Benign false refusal: 49.3% (low is good)
  • Benign helpful completion: 55.6% (high is good)
  • IFEval: 48.4%
  • GSM8K: 56.8%
  • MMLU: 50.0%

Training involved 40,000 SFT examples (general, harmful-refusal, benign-helpfulness) and 3,997 DPO pairs focusing on harmful safety preference, benign compliance, and general helpfulness. It utilizes LoRA with rank 32 and alpha 64.

Intended Use Cases

This model is primarily intended for:

  • Research on refusal alignment, safety–helpfulness trade-offs, and mechanistic interpretability.
  • Reproduction of the associated behavioral evaluation and direction-ablation studies.

It is explicitly not intended for high-stakes, safety-critical, or production deployment, nor as a replacement for application-specific safety policies. Users should note that it retains substantial benign false-refusal and capability loss, making it suitable for research purposes only.