sirigidipandu/qwen3-4b-instruct-2507

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 11, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The Qwen3-4B-Instruct-2507 is a 4.0 billion parameter instruction-tuned causal language model developed by Qwen, featuring significant enhancements in general capabilities including instruction following, logical reasoning, and coding. This updated version, with a native context length of 262,144 tokens, excels in long-tail knowledge coverage across multiple languages and offers improved alignment with user preferences for subjective and open-ended tasks. It is particularly optimized for complex reasoning, mathematics, and agentic use cases, demonstrating strong performance across various benchmarks.

Loading preview...

Qwen3-4B-Instruct-2507: Enhanced Instruction-Tuned Model

Qwen3-4B-Instruct-2507 is an updated 4.0 billion parameter causal language model from Qwen, building upon the Qwen3-4B Uncensored mode. This version introduces substantial improvements across a range of capabilities, making it a versatile choice for various applications.

Key Enhancements and Capabilities

  • General Capabilities: Significant gains in instruction following, logical reasoning, text comprehension, mathematics, science, coding, and tool usage.
  • Long-Tail Knowledge: Substantial improvements in knowledge coverage across multiple languages.
  • User Alignment: Markedly better alignment with user preferences for subjective and open-ended tasks, leading to more helpful and higher-quality text generation.
  • Extended Context Window: Features an impressive native context length of 262,144 tokens, enhancing its ability to understand and process long inputs.
  • Agentic Use: Excels in tool-calling capabilities, with recommendations to use Qwen-Agent for optimal performance.

Performance Highlights

Benchmarking against other models, Qwen3-4B-Instruct-2507 demonstrates strong performance:

  • Knowledge: Achieves 69.6 on MMLU-Pro and 84.2 on MMLU-Redux, outperforming its predecessor and other models.
  • Reasoning: Scores 47.4 on AIME25, 31.0 on HMMT25, and 80.2 on ZebraLogic, indicating strong logical reasoning abilities.
  • Coding: Attains 35.1 on LiveCodeBench v6 and 76.8 on MultiPL-E.
  • Alignment: Shows high scores in IFEval (83.4), Arena-Hard v2 (43.4), Creative Writing v3 (83.5), and WritingBench (83.4).

Recommended Use Cases

  • Complex Instruction Following: Ideal for tasks requiring precise adherence to instructions.
  • Advanced Reasoning: Suitable for mathematical, scientific, and logical problem-solving.
  • Long-Context Applications: Excellent for processing and generating content based on very long inputs.
  • Multilingual Applications: Benefits from enhanced long-tail knowledge coverage across various languages.
  • Agent-based Systems: Highly effective for tool-calling and integration into agentic workflows.