prithivMLmods/Qwen3.5-2B-Opus-Distilled-Heretic-Thinking-Multistage-SFT-v1.0

VISIONConcurrent Unit Cost:1Model Size:2.3BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 7, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

prithivMLmods/Qwen3.5-2B-Opus-Distilled-Heretic-Thinking-Multistage-SFT-v1.0 is a 2.3 billion parameter language model built on Qwen/Qwen3.5-2B, fine-tuned using a multi-stage supervised fine-tuning pipeline. It was trained on approximately 6,000 coding and STEM-focused Opus reasoning traces to enhance long-form reasoning, coding, mathematical problem-solving, and scientific analysis. This model specializes in complex reasoning tasks and instruction following, making it suitable for research and efficient local deployment in STEM and coding environments. It supports a long context length of 32,768 tokens.

Loading preview...

Model Overview

prithivMLmods/Qwen3.5-2B-Opus-Distilled-Heretic-Thinking-Multistage-SFT-v1.0 is a 2.3 billion parameter language model developed by prithivMLmods, based on the Qwen/Qwen3.5-2B architecture. This model is specifically designed to enhance reasoning capabilities across various domains, including coding, mathematics, and scientific analysis.

Key Capabilities and Training

  • Foundation: Built upon the robust Qwen 3.5-2B base model.
  • Multi-Stage SFT: Utilizes a multi-stage supervised fine-tuning pipeline for progressive improvement in reasoning performance.
  • Opus Reasoning Distillation: A core differentiator is its training on approximately 6,000 coding and STEM-focused Opus reasoning traces, alongside other high-quality reasoning data.
  • Enhanced Reasoning: Focuses on strengthening long-form reasoning, multi-step analytical capabilities, and instruction following.
  • Long Context: Supports a maximum sequence length of 32,768 tokens.
  • Efficient Deployment: As a 2B-parameter model, it is optimized for local inference and research environments.

Intended Use Cases

  • Reasoning Research: Ideal for studying reasoning distillation and multi-stage SFT techniques.
  • Coding Assistance: Improves code understanding, generation, and debugging through distilled reasoning.
  • STEM Problem Solving: Excels at mathematical, scientific, and engineering problems requiring structured reasoning.
  • Instruction Following: Useful for evaluating and enhancing multi-step instruction-following capabilities.

Limitations

As an experimental release, the model may exhibit unexpected behaviors or reasoning artifacts. Complex reasoning chains might occasionally produce incorrect intermediate steps or conclusions, and its performance reflects the biases of its training data.