SpectreFestival/Qwen2.5-3B-Instruct

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 26, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

SpectreFestival/Qwen2.5-3B-Instruct is a 3.09 billion parameter instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen. It features a 32,768 token context length and excels in instruction following, long text generation, structured data understanding, and generating structured outputs like JSON. This model demonstrates significant improvements in coding and mathematics capabilities, making it suitable for a wide range of advanced NLP tasks.

Loading preview...

Qwen2.5-3B-Instruct Overview

Qwen2.5-3B-Instruct is an instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen. This 3.09 billion parameter model builds upon its predecessors with substantial enhancements across several key areas. It supports a full context length of 32,768 tokens and can generate up to 8,192 tokens.

Key Capabilities

  • Enhanced Knowledge & Reasoning: Significantly improved capabilities in coding and mathematics, leveraging specialized expert models.
  • Instruction Following: Demonstrates strong instruction following, even with diverse system prompts, improving role-play and chatbot condition-setting.
  • Long Text Generation: Excels at generating long texts, supporting outputs over 8,000 tokens.
  • Structured Data & Output: Improved understanding of structured data (e.g., tables) and generation of structured outputs, particularly JSON.
  • Multilingual Support: Offers robust support for over 29 languages, including Chinese, English, French, Spanish, German, and Japanese.

Architecture & Training

The model utilizes a transformer architecture with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings. It has 36 layers and 16 attention heads for Q, with 2 for KV (GQA). The model underwent both pretraining and post-training stages to achieve its refined performance.

Good For

  • Applications requiring strong instruction adherence and complex task execution.
  • Code generation and mathematical problem-solving.
  • Generating lengthy and coherent textual content.
  • Processing and generating structured data, such as JSON outputs.
  • Multilingual applications and chatbots requiring robust system prompt resilience.