Mahesh111000/qwen3-8b-hanabi-grpo-step_210

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 1, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Mahesh111000/qwen3-8b-hanabi-grpo-step_210 is an 8.2 billion parameter causal language model from the Qwen3 series, developed by Qwen. This model uniquely supports seamless switching between a 'thinking mode' for complex logical reasoning, math, and coding, and a 'non-thinking mode' for efficient general-purpose dialogue. It features enhanced reasoning capabilities, superior human preference alignment for creative writing and role-playing, and strong agent capabilities with tool integration. The model supports a native context length of 32,768 tokens, extendable to 131,072 tokens using YaRN, and offers multilingual instruction following across 100+ languages.

Loading preview...

Qwen3-8B: A Versatile Language Model with Dynamic Reasoning

Mahesh111000/qwen3-8b-hanabi-grpo-step_210 is an 8.2 billion parameter model from the Qwen3 series, designed to offer advanced capabilities in reasoning, instruction-following, and agent tasks. A key differentiator is its unique ability to switch between a 'thinking mode' and a 'non-thinking mode' within a single model, optimizing performance for diverse scenarios.

Key Capabilities

  • Dynamic Reasoning: Seamlessly transitions between a detailed 'thinking mode' for complex logical reasoning, mathematics, and code generation, and an efficient 'non-thinking mode' for general dialogue. This enhances performance beyond previous Qwen models in both contexts.
  • Human Preference Alignment: Excels in creative writing, role-playing, multi-turn conversations, and instruction following, providing a more natural and engaging user experience.
  • Agentic Functionality: Demonstrates strong tool-calling capabilities, integrating precisely with external tools in both thinking and unthinking modes, achieving leading performance in complex agent-based tasks among open-source models.
  • Multilingual Support: Supports over 100 languages and dialects, offering robust multilingual instruction following and translation.
  • Extended Context Window: Natively handles up to 32,768 tokens, with validated support for up to 131,072 tokens using the YaRN method for long text processing.

Good for

  • Applications requiring adaptive reasoning, where the model needs to perform both complex problem-solving and efficient general conversation.
  • Creative content generation and role-playing scenarios due to its superior human preference alignment.
  • Agent-based systems and tool-use applications that benefit from precise integration and strong performance.
  • Multilingual applications needing robust instruction following and translation across many languages.
  • Use cases involving long documents or conversations that require extended context lengths.