roskosmos19/Orca-4B-Thinking
Whale-4B-Thinking by roskosmos19 is a 4 billion parameter causal language model based on the Qwen3-4B architecture, specialized for extreme Chain-of-Thought (CoT) reasoning. It features a native 262,144 token context length and is optimized for deep, self-reflective thinking processes in the DeepSeek-R1 style. This model excels in complex reasoning tasks across mathematics, science, and code, demonstrating significant performance gains on benchmarks like AIME, GPQA, and MATH. It is designed for applications requiring extensive, structured thought processes rather than direct answers.
Loading preview...
Whale-4B-Thinking: Extreme CoT Reasoning
Whale-4B-Thinking is a 4 billion parameter language model developed by roskosmos19, built upon the Qwen3-4B architecture. Its core specialization is extreme Chain-of-Thought (CoT) reasoning, designed to emulate the deep, self-reflective thinking style of DeepSeek-R1. This model is fine-tuned using high-quality DeepSeek-R1 distilled datasets, specifically "open-thoughts/OpenThoughts-114k" and "open-r1/Mixture-of-Thoughts", which contain verified reasoning traces across math, code, science, and puzzles.
Key Capabilities
- Deep Reasoning: Generates long, structured Chain-of-Thoughts with self-verification, reflection, and deep exploration, directly in the DeepSeek-R1 style.
- Enhanced Performance: Shows significant improvements on reasoning-intensive benchmarks such as AIME (81–83%), GPQA (65–67%), LiveCodeBench (55–57%), and MATH (90%+).
- Ultra-Long Context: Features a native 262,144 token context length, enabling the processing of very extensive reasoning traces.
- Instruction Following & Tool Use: Improved general capabilities in instruction following, tool use, and alignment, while maintaining its dedicated thinking mode.
When to Use This Model
Whale-4B-Thinking is ideal for use cases demanding profound, step-by-step problem-solving and complex analytical tasks. It operates exclusively in a "Thinking-Mode," where the model first generates a detailed thought process before providing a final answer. This makes it particularly suitable for:
- Mathematical Problem Solving: Excelling in complex math problems requiring detailed derivations.
- Code Generation & Debugging: Generating logical steps for coding challenges.
- Scientific Inquiry: Tackling scientific questions that benefit from structured reasoning.
- Any task requiring transparent, verifiable thought processes.
It is recommended to use specific sampling parameters (Temperature=0.6, TopP=0.95, TopK=20) and allow for long output lengths (up to 81,920 tokens) to fully leverage its reasoning capabilities.