tangyunbo/RL

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 6, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen3-4B-Instruct-2507 is a 4 billion parameter causal language model developed by Qwen, featuring significant enhancements in general capabilities including instruction following, logical reasoning, and coding. This updated version of the Qwen3-4B non-thinking mode excels in long-tail knowledge coverage across multiple languages and offers enhanced 256K long-context understanding. It is particularly strong in subjective and open-ended tasks, providing more helpful responses and higher-quality text generation.

Loading preview...

Qwen3-4B-Instruct-2507: Enhanced 4B Causal Language Model

Qwen3-4B-Instruct-2507 is an updated 4 billion parameter causal language model from Qwen, building upon the Qwen3-4B non-thinking mode. It features a native context length of 262,144 tokens, making it highly capable for long-context understanding tasks. This model is designed to operate without generating <think></think> blocks, simplifying its output for direct instruction following.

Key Capabilities & Enhancements

  • General Capabilities: Significant improvements across instruction following, logical reasoning, text comprehension, mathematics, science, coding, and tool usage.
  • Knowledge & Multilingualism: Substantial gains in long-tail knowledge coverage and strong performance in multilingual benchmarks like MultiIF and PolyMATH.
  • Alignment & Subjectivity: Markedly better alignment with user preferences in subjective and open-ended tasks, leading to more helpful and higher-quality text generation.
  • Agentic Use: Excels in tool calling capabilities, with recommendations to use Qwen-Agent for optimal performance.

Performance Highlights

Benchmarking against previous versions and other models, Qwen3-4B-Instruct-2507 shows notable improvements:

  • Knowledge: Achieves 69.6 on MMLU-Pro and 84.2 on MMLU-Redux.
  • Reasoning: Scores 47.4 on AIME25 and 80.2 on ZebraLogic.
  • Coding: Reaches 35.1 on LiveCodeBench v6 and 76.8 on MultiPL-E.
  • Alignment: Demonstrates strong performance in Creative Writing v3 (83.5) and WritingBench (83.4).

Recommended Use Cases

This model is well-suited for applications requiring:

  • Advanced instruction following and complex reasoning.
  • Processing and generating long-form content due to its extended context window.
  • Multilingual text generation and comprehension.
  • Agentic workflows and tool-use scenarios.
  • Tasks demanding high-quality, aligned, and subjective text generation.