roskosmos19/Whale-4B-Thinking-2507

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 4, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The roskosmos19/Whale-4B-Thinking-2507 is a 4 billion parameter causal language model based on the Qwen3-4B architecture, specialized for extreme Chain-of-Thought (CoT) reasoning. It features a native 262,144 token context length and is optimized for deep, self-reflective reasoning traces in the DeepSeek-R1 style. This model excels in complex problem-solving across mathematics, science, and code, demonstrating significant performance gains on reasoning-intensive benchmarks.

Loading preview...

Whale-4B-Thinking: Extreme Reasoning Model

Whale-4B-Thinking, developed by roskosmos19, is a 4 billion parameter language model built upon the Qwen3-4B architecture, specifically engineered for advanced Chain-of-Thought (CoT) reasoning. It distinguishes itself through its ability to generate long, structured, and self-correcting thought processes, mirroring the style of DeepSeek-R1.

Key Capabilities and Features

  • Extreme CoT Reasoning: Generates extensive, structured reasoning traces with self-verification, reflection, and deep exploration, directly inspired by DeepSeek-R1.
  • Enhanced Reasoning Performance: Shows substantial improvements on challenging benchmarks such as AIME, GPQA, LiveCodeBench, and MATH.
  • Improved General Abilities: Offers better instruction-following, tool-use, and alignment while operating in a dedicated "Thinking-Mode."
  • Ultra-Long Context: Supports a native context length of 262,144 tokens, ideal for very long reasoning sequences.
  • Specialized Training: Fine-tuned on high-quality DeepSeek-R1 distilled datasets, including open-thoughts/OpenThoughts-114k and open-r1/Mixture-of-Thoughts, which contain verified reasoning traces across multiple domains.

When to Use This Model

Whale-4B-Thinking is particularly well-suited for applications requiring deep, step-by-step problem-solving and complex logical deduction. It operates exclusively in a "Thinking-Mode," where the chat template automatically injects the reasoning process. This model is recommended for tasks in:

  • Mathematics and Science: Solving intricate problems that benefit from detailed, verifiable reasoning steps.
  • Code Generation and Analysis: Tackling complex coding challenges with structured thought processes.
  • Any domain requiring robust, self-reflective reasoning: Where a transparent and verifiable chain of thought is crucial for accuracy and understanding.