tri-fair-lab/Snowdon1.1-Small
Snowdon1.1-Small is a 35.1 billion parameter causal language model developed by tri-fair-lab, re-aligned from Qwen3.6-35B-A3B. This model is specifically optimized to reduce topic-conditioned misalignment and provide balanced framing on politically sensitive subjects, while preserving general capabilities. It achieves this through a novel two-stage re-alignment pipeline involving Fisher-routed directional ablation and Constitutional DPO, making it suitable for applications requiring impartial and constitution-aligned responses.
Loading preview...
Snowdon1.1-Small: Constitutionally Re-aligned LLM
Snowdon1.1-Small, developed by tri-fair-lab, is a 35.1 billion parameter causal language model based on Qwen3.6-35B-A3B. Its core innovation lies in its re-alignment to the Public AI Constitution, explicitly targeting and reducing "topic-conditioned misalignment." This means the model is designed to engage with the substance of queries, regardless of political sensitivity or region, and present contested subjects with balanced framing, rather than exhibiting refusal or biased perspectives.
Key Capabilities & Differentiators
- Explicit, Auditable Alignment: The model's behavior is governed by a written constitution, allowing for transparent evaluation and revision of its ethical and impartiality standards.
- Capability Preservation: The re-alignment process, utilizing Fisher-routed directional ablation and Constitutional DPO, is engineered to modify the model's values without significantly degrading its general capabilities. Benchmarks show minimal impact on general tasks like MMLU-Pro and GPQA-Diamond, while drastically improving re-alignment scores (e.g., 93.0 on Perspective Bench vs. 16.0 for the base model).
- Safety Refusal Maintained: Unlike some re-alignment methods, Snowdon1.1-Small preserves safety refusal, maintaining 97.8% refusal on unsafe requests, comparable to the base model's 98.5%.
- Efficient Re-alignment: The re-alignment is achieved through a merged weight edit, meaning no added parameters or inference-time latency, making it cost-effective and efficient for deployment.
Ideal Use Cases
- Applications requiring impartial responses: Suitable for scenarios where neutrality and balanced perspectives on sensitive topics are critical.
- Ethical AI development: For developers building systems that need to adhere to a transparent and auditable set of constitutional values.
- Content moderation and analysis: Can be used to analyze or generate content that avoids topic-conditioned bias or refusal.