C10X/Qwen2.5-1.5B-Instruct

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 10, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

C10X/Qwen2.5-1.5B-Instruct is a 1.54 billion parameter instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen. It features a 32,768 token context length and is significantly improved in coding, mathematics, and instruction following. This model excels at generating long texts, understanding structured data, and producing structured outputs like JSON, while also offering robust multilingual support for over 29 languages.

Loading preview...

Qwen2.5-1.5B-Instruct Overview

Qwen2.5-1.5B-Instruct is an instruction-tuned causal language model from the Qwen2.5 series, featuring 1.54 billion parameters and a substantial 32,768 token context length. Developed by Qwen, this model builds upon its predecessors with significant enhancements across several key areas.

Key Capabilities & Improvements

  • Enhanced Knowledge & Specialized Skills: Demonstrates greatly improved capabilities in coding and mathematics, benefiting from specialized expert models.
  • Instruction Following: Shows significant improvements in adhering to instructions and handling diverse system prompts, making it more resilient for chatbot implementations and role-play scenarios.
  • Text Generation & Understanding: Excels at generating long texts (up to 8K tokens), understanding structured data (e.g., tables), and generating structured outputs, particularly JSON.
  • Multilingual Support: Offers robust support for over 29 languages, including major global languages like Chinese, English, French, Spanish, German, and Japanese.
  • Long-Context Processing: Supports a full context length of 32,768 tokens, with the ability to generate up to 8,192 tokens.

Architecture Highlights

The model utilizes a transformer architecture incorporating RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings. It consists of 28 layers and employs Grouped-query Attention (GQA) with 12 heads for Q and 2 for KV.

When to Use This Model

This model is particularly well-suited for applications requiring:

  • Code generation and mathematical problem-solving.
  • Complex instruction following and structured output generation (e.g., JSON).
  • Long-form content generation and processing extensive textual inputs.
  • Multilingual applications across a broad range of languages.