LIF1014/ptdbench-verl-coding-tasks-function-call

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 22, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen2.5-1.5B-Instruct is a 1.54 billion parameter instruction-tuned causal language model from the Qwen2.5 series by Qwen Team. This model significantly improves capabilities in coding and mathematics, instruction following, and generating structured outputs like JSON. It supports a full context length of 32,768 tokens and is multilingual, supporting over 29 languages. It is optimized for diverse chatbot implementations and handling complex system prompts.

Loading preview...

Qwen2.5-1.5B-Instruct Overview

Qwen2.5-1.5B-Instruct is a 1.54 billion parameter instruction-tuned causal language model, part of the Qwen2.5 series developed by the Qwen Team. This model builds upon its predecessors with significant enhancements across several key areas, making it a versatile tool for various NLP applications.

Key Capabilities & Improvements

  • Enhanced Knowledge & Reasoning: Demonstrates significantly improved capabilities in coding and mathematics, leveraging specialized expert models.
  • Instruction Following: Shows substantial improvements in adhering to instructions and generating structured outputs, particularly JSON.
  • Long Text Generation: Excels at generating long texts, supporting outputs of up to 8,000 tokens.
  • Context Length: Features a robust 32,768-token context window for processing extensive inputs.
  • Multilingual Support: Offers broad multilingual capabilities, supporting over 29 languages including Chinese, English, French, Spanish, and more.
  • Robustness: More resilient to diverse system prompts, enhancing its performance in role-play and complex chatbot scenarios.

Architecture & Features

This model utilizes a transformer architecture with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings. It has 28 layers and 12 attention heads (with 2 for KV in GQA configuration). The model is designed for both pretraining and post-training stages, focusing on delivering high performance in instruction-tuned tasks.

Ideal Use Cases

This model is particularly well-suited for applications requiring strong coding assistance, mathematical problem-solving, precise instruction following, generation of structured data, and multilingual communication within a substantial context window.