Nexa-AI-Official/Nexa-AI-4B-Instruct
Nexa-AI-4B-Instruct is a 4 billion parameter instruction-tuned large language model developed by Neura Tech AI and Lumina AI, built upon Qwen/Qwen3-4B-Instruct-2507. This model features a 262,144 token context length and excels in general conversation, instruction following, coding, mathematics, logical reasoning, and multilingual understanding. It is particularly optimized for agent and tool calling support, demonstrating strong performance across various benchmarks in reasoning, coding, and agentic tasks.
Loading preview...
Nexa-AI-4B-Instruct: An Overview
Nexa-AI-4B-Instruct is a 4 billion parameter instruction-tuned large language model, a collaborative effort by Neura Tech AI and Lumina AI. It is based on the Qwen/Qwen3-4B-Instruct-2507 model and inherits its substantial 262,144 token context length. This model is designed as a capable multilingual AI assistant, focusing on a broad range of applications.
Key Capabilities and Features
- Instruction Following: High-quality adherence to user instructions.
- Multilingual Support: Strong understanding across multiple languages, including English, Hindi, and Chinese.
- Reasoning & Mathematics: Enhanced capabilities in logical reasoning and mathematical problem-solving, outperforming its base model and other comparably sized models on benchmarks like AIME25, HMMT25, and ZebraLogic.
- Coding Assistance: Provides robust support for coding tasks, showing competitive performance on LiveCodeBench and MultiPL-E.
- Agent & Tool Calling: Specifically optimized for AI agent workflows and tool integration, demonstrating superior results in BFCL-v3 and TAU benchmarks.
- Long Context Understanding: Benefits from the extensive context window of its base architecture.
Performance Highlights
Nexa-AI-4B-Instruct shows significant improvements over its base model, Qwen3-4B, and often surpasses larger models like Qwen3-30B in specific areas. Notable benchmark wins include:
- MMLU-Pro: 69.6%
- GPQA: 62.0%
- AIME25: 47.4%
- ZebraLogic: 80.2%
- LiveCodeBench v6: 35.1%
- Arena-Hard v2: 43.4%
- Creative Writing v3: 83.5%
- BFCL-v3: 61.9%
This model is particularly well-suited for applications requiring strong reasoning, coding, and agentic capabilities within a multilingual context.