sfewf/qwen3-4b-math-RL
The sfewf/qwen3-4b-math-RL model is a Qwen3-4b variant specifically post-RL trained on a Math dataset, featuring a 'Max-Thinking mode' for enhanced reasoning. This model is optimized for mathematical problem-solving and complex reasoning tasks, demonstrating improved efficiency in both standard and max-thinking modes. It excels in accuracy on benchmarks like GSM8K, MATH-lighteval, BBH, and GPQA, making it suitable for applications requiring robust analytical capabilities.
Loading preview...
Model Overview
This model, sfewf/qwen3-4b-math-RL, is a specialized variant of the Qwen3-4b architecture that has undergone Reinforcement Learning (RL) training on a comprehensive Math dataset. A key feature is its Max-Thinking mode, inspired by DeepSeek V4, which encourages thorough, step-by-step deliberation for complex problems.
Key Capabilities
- Enhanced Mathematical Reasoning: Demonstrates significant improvements in solving mathematical problems, as evidenced by benchmark scores.
- Max-Thinking Mode: Provides a detailed, rigorous deliberation process, documenting intermediate steps and considered alternatives.
- Efficient Reasoning: Observations indicate more efficient reasoning in both default and Max-Thinking modes.
- Strong Benchmark Performance: Achieves high accuracy on various datasets:
- GSM8K: 0.9172 (standard), 0.9327 (max-effort)
- MATH-lighteval: 0.8019 (standard), 0.8505 (max-effort)
- BBH: 0.7963 (standard), 0.8709 (max-effort)
- GPQA: 0.2667 (standard), 0.3125 (max-effort)
- Length-Penalty Training: Responses are concise while maintaining accuracy in default mode.
Good For
- Mathematical Problem Solving: Ideal for applications requiring precise and detailed solutions to math-related queries.
- Complex Reasoning Tasks: Benefits from the Max-Thinking mode for problems demanding extensive logical decomposition and verification.
- Educational Tools: Can be used in systems that require step-by-step explanations for mathematical concepts.
- Research and Development: Suitable for exploring advanced reasoning capabilities in LLMs.