prithivMLmods/Q3.5-9B-DS-v4-Flash-DA
prithivMLmods/Q3.5-9B-DS-v4-Flash-DA is a 9 billion parameter language model built on the Qwen3.5 backbone, specifically prithivMLmods/Qwen3.5-9B-Unredacted-MAX. It is optimized for rich, detailed, and context-aware reasoning through multi-stage distillation on DeepSeek V4 reasoning traces. This model utilizes advanced refusal direction analysis and ablation-based training to reduce internal refusal behaviors while preserving strong reasoning and instruction-following performance, making it suitable for complex reasoning and red-teaming research.
Loading preview...
Overview
Q3.5-9B-DS-v4-Flash-DA is a 9 billion parameter model developed by prithivMLmods, based on the Qwen/Qwen3.5-9B architecture and further refined from prithivMLmods/Qwen3.5-9B-Unredacted-MAX. Its core innovation lies in its "Distilled-Abliterated" approach, which involves multi-stage distillation using curated reasoning traces from DeepSeek V4 Flash. This process significantly enhances its multi-step reasoning capabilities.
Key Capabilities
- Enhanced Reasoning: Fine-tuned with DeepSeek V4 reasoning traces for improved multi-step and long-chain thinking.
- Reduced Refusal Behaviors: Employs advanced refusal direction analysis and ablation strategies to minimize internal refusals while maintaining reasoning quality.
- Instruction Following: Seamlessly handles both complex reasoning and instruction adherence.
- High-Coherence Outputs: Designed to produce consistent and contextually grounded outputs, even in long generations.
Intended Use Cases
- Reasoning & Chain-of-Thought Tasks: Excels in deep, multi-step reasoning scenarios.
- Instruction Following: Ideal for hybrid prompts requiring both instruction adherence and complex reasoning.
- Red-Teaming & Alignment Research: Useful for evaluating systems with reduced refusal mechanisms and studying refusal direction analysis.
- Research on Abliteration: Provides a platform for studying the effects of ablation-based training on reasoning preservation.
Limitations
It is important to note that this model intentionally minimizes built-in safety refusals, which means it may generate sensitive or unrestricted content. Users are responsible for ethical usage, and the model may require significant VRAM for deployment.