Fergusfang/Qwen2.5-3B-Instruct

TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 13, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

Fergusfang/Qwen2.5-3B-Instruct is a 3.09 billion parameter instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen. It features a 32,768 token context length and is built on a transformer architecture with RoPE, SwiGLU, and RMSNorm. This model significantly enhances capabilities in coding, mathematics, instruction following, and generating structured outputs like JSON, while also supporting over 29 languages.

Loading preview...

Qwen2.5-3B-Instruct Overview

Fergusfang/Qwen2.5-3B-Instruct is a 3.09 billion parameter instruction-tuned causal language model, part of the Qwen2.5 series developed by Qwen. This model builds upon its predecessor, Qwen2, with notable improvements across several key areas. It supports a substantial context length of 32,768 tokens for input and can generate up to 8,192 tokens.

Key Capabilities & Improvements

  • Enhanced Knowledge & Reasoning: Significantly improved capabilities in coding and mathematics, leveraging specialized expert models.
  • Instruction Following: Demonstrates substantial advancements in adhering to instructions and generating long-form text (over 8K tokens).
  • Structured Data & Output: Better understanding of structured data, such as tables, and improved generation of structured outputs, particularly JSON.
  • Robustness: More resilient to diverse system prompts, which benefits role-play implementations and chatbot condition-setting.
  • Multilingual Support: Offers support for over 29 languages, including major global languages like Chinese, English, French, Spanish, and Japanese.

Architecture & Technical Specifications

This model utilizes a transformer architecture incorporating RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings. It features 36 layers and 16 attention heads for Q with 2 for KV (GQA). The model is designed for both pretraining and post-training stages, making it a versatile foundation for various NLP tasks.