reaperdoesntknow/DiStil-Qwen3-1.7B-uncensored
DiStil-Qwen3-1.7B-uncensored is a 1.7 billion effective parameter Qwen3ForCausalLM model developed by Convergent Intelligence LLC: Research Division. It is produced by distilling Qwen3 with uncensored SFT data, specifically designed to remove alignment-imposed refusal behaviors while preserving the base model's reasoning and generation capabilities. This model features a 40,960 token context length and aims to respond directly to prompts without filtering through safety heuristics, making it suitable for technical, analytical, and research queries requiring unfiltered responses.
Loading preview...
Model Overview
DiStil-Qwen3-1.7B-uncensored is a 1.7 billion effective parameter model from the Qwen3 family, developed by Convergent Intelligence LLC: Research Division. It is a distilled version of Qwen3, specifically fine-tuned using uncensored SFT data. The primary goal of this model is to eliminate alignment-imposed refusal behaviors, ensuring it responds directly to prompts without filtering through safety heuristics that might interfere with legitimate technical, analytical, or research inquiries.
Key Characteristics
- Architecture: Based on Qwen3ForCausalLM, maintaining the original architecture and tokenizer.
- Parameters: Approximately 2.03 billion total parameters, with an effective 1.7 billion.
- Context Length: Supports a substantial context window of 40,960 tokens.
- Training: Supervised fine-tuning (SFT) using TRL on uncensored instruction data, focusing on shifting the model's response distribution away from refusal patterns without architectural modifications.
- Uncensored Nature: Designed to provide unfiltered responses, preserving the base model's reasoning and generation capabilities without safety-related interference.
Unique Approach
This model is part of a distillation chain built upon Discrepancy Calculus (DISC), a measure-theoretic framework for analyzing and quantifying local structural mismatches in model outputs. This advanced methodology aims to transfer capabilities effectively while removing unwanted behaviors. The full methodology is detailed in "Structure Over Scale" (DOI: 10.57967/hf/8165).
Ideal Use Cases
- Applications requiring direct, unfiltered responses to complex or sensitive queries.
- Research and development where alignment-imposed refusals hinder analytical tasks.
- Scenarios where preserving the base model's raw reasoning capabilities is paramount.