prithivMLmods/Qwen3.5-4B-Opus-Distilled-Heretic-Thinking-Multistage-SFT-v1.0
prithivMLmods/Qwen3.5-4B-Opus-Distilled-Heretic-Thinking-Multistage-SFT-v1.0 is a 4.5 billion parameter language model built on Qwen/Qwen3.5-4B, fine-tuned using a multi-stage supervised fine-tuning (SFT) pipeline. It was distilled from approximately 6,000 coding and STEM-focused Opus reasoning traces to enhance long-form reasoning, coding, mathematical problem-solving, and scientific analysis. This model is optimized for research and experimentation in complex reasoning tasks and efficient local deployment.
Loading preview...
Model Overview
Qwen3.5-4B-Opus-Distilled-Heretic-Thinking-Multistage-SFT-v1.0 is a 4.5 billion parameter language model developed by prithivMLmods, based on the Qwen/Qwen3.5-4B architecture. This model is specifically designed to excel in complex reasoning tasks, achieved through a unique multi-stage supervised fine-tuning (SFT) process. It leverages distillation from approximately 6,000 high-quality coding and STEM-focused Opus reasoning traces, alongside additional reasoning data, to significantly improve its analytical and problem-solving capabilities.
Key Capabilities
- Enhanced Reasoning: Specialized in long-form reasoning, multi-step analytical tasks, and instruction following.
- STEM and Coding Proficiency: Optimized for mathematical problem-solving, scientific analysis, and coding assistance, including code understanding and generation.
- Multi-Stage SFT: Utilizes a progressive fine-tuning pipeline for continuous performance improvement in reasoning.
- Efficient Deployment: As a 4.5B parameter model, it is suitable for local inference and research environments.
- Research Focus: Primarily intended for research and experimentation into reasoning distillation techniques.
Intended Use Cases
- Reasoning Research: Ideal for studying advanced reasoning distillation and multi-stage SFT methods.
- Coding Assistance: Useful for tasks requiring code understanding, generation, and debugging.
- STEM Problem Solving: Applicable to solving complex problems in mathematics, science, and engineering.
- Instruction Following: Evaluating and improving the model's ability to follow multi-step instructions.
This model is released as an experimental version, and users should be aware of potential reasoning artifacts or unexpected behaviors in certain complex scenarios.