sirigidipandu/qwen3-4b-instruct-2507
The Qwen3-4B-Instruct-2507 is a 4.0 billion parameter instruction-tuned causal language model developed by Qwen, featuring significant enhancements in general capabilities including instruction following, logical reasoning, and coding. This updated version, with a native context length of 262,144 tokens, excels in long-tail knowledge coverage across multiple languages and offers improved alignment with user preferences for subjective and open-ended tasks. It is particularly optimized for complex reasoning, mathematics, and agentic use cases, demonstrating strong performance across various benchmarks.
Loading preview...
Qwen3-4B-Instruct-2507: Enhanced Instruction-Tuned Model
Qwen3-4B-Instruct-2507 is an updated 4.0 billion parameter causal language model from Qwen, building upon the Qwen3-4B Uncensored mode. This version introduces substantial improvements across a range of capabilities, making it a versatile choice for various applications.
Key Enhancements and Capabilities
- General Capabilities: Significant gains in instruction following, logical reasoning, text comprehension, mathematics, science, coding, and tool usage.
- Long-Tail Knowledge: Substantial improvements in knowledge coverage across multiple languages.
- User Alignment: Markedly better alignment with user preferences for subjective and open-ended tasks, leading to more helpful and higher-quality text generation.
- Extended Context Window: Features an impressive native context length of 262,144 tokens, enhancing its ability to understand and process long inputs.
- Agentic Use: Excels in tool-calling capabilities, with recommendations to use Qwen-Agent for optimal performance.
Performance Highlights
Benchmarking against other models, Qwen3-4B-Instruct-2507 demonstrates strong performance:
- Knowledge: Achieves 69.6 on MMLU-Pro and 84.2 on MMLU-Redux, outperforming its predecessor and other models.
- Reasoning: Scores 47.4 on AIME25, 31.0 on HMMT25, and 80.2 on ZebraLogic, indicating strong logical reasoning abilities.
- Coding: Attains 35.1 on LiveCodeBench v6 and 76.8 on MultiPL-E.
- Alignment: Shows high scores in IFEval (83.4), Arena-Hard v2 (43.4), Creative Writing v3 (83.5), and WritingBench (83.4).
Recommended Use Cases
- Complex Instruction Following: Ideal for tasks requiring precise adherence to instructions.
- Advanced Reasoning: Suitable for mathematical, scientific, and logical problem-solving.
- Long-Context Applications: Excellent for processing and generating content based on very long inputs.
- Multilingual Applications: Benefits from enhanced long-tail knowledge coverage across various languages.
- Agent-based Systems: Highly effective for tool-calling and integration into agentic workflows.