TreezzZ/r2s-alfworld-qwen3-8b-student
TreezzZ/r2s-alfworld-qwen3-8b-student is an 8.2 billion parameter causal language model from the Qwen3 series, developed by Qwen. It uniquely supports seamless switching between a 'thinking mode' for complex reasoning (math, code) and a 'non-thinking mode' for efficient general dialogue. This model excels in reasoning capabilities, human preference alignment, and agentic tasks, supporting over 100 languages with a native context length of 32,768 tokens, extendable to 131,072 tokens with YaRN.
Loading preview...
Qwen3-8B: A Versatile Language Model with Dynamic Thinking Modes
Qwen3-8B is an 8.2 billion parameter causal language model from the Qwen series, designed for advanced reasoning, instruction-following, and agent capabilities. It introduces a novel feature allowing seamless switching between two operational modes:
Key Capabilities
- Dynamic Thinking Modes: Uniquely supports a 'thinking mode' for complex logical reasoning, mathematics, and coding, and a 'non-thinking mode' for efficient, general-purpose dialogue within a single model. This ensures optimal performance across diverse scenarios.
- Enhanced Reasoning: Demonstrates significant improvements in mathematical problem-solving, code generation, and commonsense logical reasoning, outperforming previous Qwen models in both thinking and non-thinking configurations.
- Superior Human Preference Alignment: Excels in creative writing, role-playing, multi-turn conversations, and instruction following, providing a more natural and engaging user experience.
- Advanced Agentic Functionality: Offers strong tool-calling capabilities, enabling precise integration with external tools in both thinking and non-thinking modes, achieving leading performance in complex agent-based tasks among open-source models.
- Multilingual Support: Supports over 100 languages and dialects, with robust capabilities for multilingual instruction following and translation.
- Extended Context Length: Natively handles up to 32,768 tokens, and can be extended to 131,072 tokens using the YaRN method for processing long texts.
Best Practices
To optimize performance, specific sampling parameters are recommended for each mode:
- Thinking Mode (
enable_thinking=True): UseTemperature=0.6,TopP=0.95,TopK=20, andMinP=0. Avoid greedy decoding. - Non-Thinking Mode (
enable_thinking=False): UseTemperature=0.7,TopP=0.8,TopK=20, andMinP=0.
For agentic use, integration with Qwen-Agent is recommended to leverage its tool-calling abilities effectively.