Dingdust/VibeThinker-3B-heretic
Dingdust/VibeThinker-3B-heretic is a 3.1 billion parameter language model, a decensored version of the VibeThinker-3B model created using the Heretic v1.4.0 tool. VibeThinker-3B, developed by WeiboAI, is specifically optimized for challenging verifiable reasoning tasks such as mathematics, coding, and STEM, achieving strong performance on benchmarks like IMO-AnswerBench and LiveCodeBench. This model demonstrates that compact models can achieve near-frontier reasoning capabilities in structured task spaces with reliable feedback signals.
Loading preview...
VibeThinker-3B-heretic: Decensored Reasoning Model
This model is a 3.1 billion parameter variant of the VibeThinker-3B, created using the Heretic v1.4.0 tool to produce a decensored version. The original VibeThinker-3B, developed by WeiboAI, is a specialized language model designed for advanced reasoning tasks, particularly in mathematics, coding, and STEM fields. It leverages the Spectrum-to-Signal Principle (SSP) post-training pipeline to achieve high performance on verifiable reasoning benchmarks.
Key Capabilities & Performance
- Verifiable Reasoning: Excels in tasks with clear verification signals, such as mathematical problem-solving, code generation, and scientific reasoning.
- Benchmark Performance: Achieves 76.4 on IMO-AnswerBench (80.6 with Claim-Level Reliability Assessment), placing it in the performance range of much larger models like DeepSeek V3.2 and GLM-5. It also passes 123/128 (96.1% acceptance rate) first-attempt submissions on recent LeetCode contests.
- Efficient Reasoning: Demonstrates that small language models (SLMs) can achieve frontier-level performance in specific, structured capability domains, challenging the notion that large parameter counts are always necessary for advanced reasoning.
Training Methodology
VibeThinker-3B follows the SSP pipeline, which includes:
- Curriculum-based two-stage SFT: Focuses on broad capability coverage initially, then shifts to harder, longer-horizon reasoning samples, using Diversity-Exploring Distillation.
- Multi-domain Reasoning RL: Applies MaxEnt-Guided Policy Optimization (MGPO) sequentially to math, code, and STEM tasks within a 64K long-context window.
- Offline Self-Distillation: Filters and distills high-quality trajectories from RL checkpoints into a unified student model.
- Instruct RL: Improves controllability for user-facing prompts using rule-based validators and rubric-based reward models.
Recommended Use Cases
- Competitive Math & Coding: Ideal for tasks requiring precise, verifiable solutions.
- STEM Reasoning: Suitable for scientific and engineering problem-solving.
- Tasks with Clear Verification: Best utilized where the correctness of the output can be objectively checked.
For broad open-domain knowledge or general-purpose dialogue, larger models may still be more appropriate. This model is particularly valuable for exploring the boundaries of small models in specific, high-performance reasoning niches.