reaperdoesntknow/TopologicalQwen
TopologicalQwen by Convergent Intelligence LLC is a 1.7 billion parameter Qwen3ForCausalLM model with a 40,960 token context length, distilled from Qwen3-30B-A3B using Topological Knowledge Distillation (TKD). This methodology captures structural information beyond standard distillation by accounting for smooth, jump, and drift components of the teacher's output distribution. It is specifically designed for complex reasoning tasks, exhibiting a DualMind format for exploration, self-critique, and synthesized responses, particularly in physics-related CoT problems.
Loading preview...
TopologicalQwen: Topology-Aware Knowledge Distillation
TopologicalQwen is a 1.7 billion parameter model developed by Convergent Intelligence LLC, distilled from a 30B Qwen3 teacher model. Its core innovation lies in Topological Knowledge Distillation (TKD), a method that goes beyond standard knowledge distillation by preserving the structural integrity of the teacher's knowledge. While conventional KD only captures smooth variations in the teacher's output, TKD explicitly accounts for three components:
- Smooth distillation: Standard continuous knowledge transfer.
- Jump corrections: Explicitly addresses discontinuities where reasoning modes or topics shift.
- Drift corrections: Captures subtle, singular-continuous distributional shifts.
This approach, grounded in Discrepancy Calculus, allows the 1.7B model to retain complex reasoning capabilities that are typically lost in smaller models. It operates with a DualMind format (<explore> → <examine> → <response>), mimicking a cognitive loop of derivation, self-critique, and synthesis, making it particularly adept at structured problem-solving like physics CoT tasks.
Key Capabilities
- Advanced Reasoning: Excels in complex, multi-step reasoning, especially in scientific domains (e.g., differential equations, theoretical mechanics).
- Structural Knowledge Preservation: TKD ensures the student model captures the 'architecture' of the teacher's knowledge, not just surface statistics.
- DualMind Cognitive Process: Generates detailed derivations, critiques its own reasoning, and then synthesizes a clean final answer.
- High Context Length: Supports a 40,960 token context, enabling processing of extensive problem descriptions and derivations.
What Makes It Different
Unlike other distillation methods, TKD leverages mathematical foundations to detect and preserve conceptual boundaries and subtle drifts in the teacher's knowledge distribution. This allows TopologicalQwen to achieve reasoning quality at 1.7B parameters that standard distillation methods cannot, even with higher compute. The model's training involved a 4-phase curriculum with topology-guided adaptive windowing and proof-weighted loss, emphasizing reasoning tokens. This model represents the application of the proven TKD methodology with premium hardware, demonstrating its effectiveness in producing structurally intelligent small models.