Mahesh111000/qwen3-8b-hanabi-grpo-step_155

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 29, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Mahesh111000/qwen3-8b-hanabi-grpo-step_155 is an 8.2 billion parameter causal language model from the Qwen3 series, developed by Qwen. This model uniquely supports seamless switching between a 'thinking mode' for complex reasoning tasks like math and coding, and a 'non-thinking mode' for general dialogue, enhancing performance across diverse scenarios. It features superior human preference alignment, strong agent capabilities, and multilingual support for over 100 languages, making it suitable for advanced conversational AI and tool-integrated applications.

Loading preview...

Model Overview

Mahesh111000/qwen3-8b-hanabi-grpo-step_155 is an 8.2 billion parameter model from the Qwen3 series, developed by Qwen. It is a causal language model that has undergone both pretraining and post-training stages. A key innovation is its ability to seamlessly switch between a 'thinking mode' for complex logical reasoning, mathematics, and coding, and a 'non-thinking mode' for efficient, general-purpose dialogue. This dual-mode functionality ensures optimal performance across various tasks.

Key Capabilities

  • Enhanced Reasoning: Significantly improved performance in mathematics, code generation, and commonsense logical reasoning compared to previous Qwen models.
  • Human Preference Alignment: Excels in creative writing, role-playing, multi-turn dialogues, and instruction following, providing a more natural conversational experience.
  • Agent Capabilities: Demonstrates strong integration with external tools in both thinking and non-thinking modes, achieving leading performance in complex agent-based tasks among open-source models.
  • Multilingual Support: Supports over 100 languages and dialects with robust multilingual instruction following and translation capabilities.
  • Extended Context: Natively supports a context length of 32,768 tokens, extendable up to 131,072 tokens using the YaRN method for processing long texts.

Good For

This model is ideal for applications requiring advanced reasoning, complex problem-solving, and sophisticated conversational interactions. Its agentic capabilities make it suitable for tool-integrated systems, while its multilingual support broadens its applicability globally. Developers can leverage its unique thinking/non-thinking mode for optimizing performance based on task complexity.