ai-team-20/kimi-dev-72b-1

TEXT GENERATIONPricing:Input $12 / Output $20Concurrent Unit Cost:4Model Size:72.7BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 11, 2025License:otherArchitecture:Transformer Featherless Exclusive Cold

Qwen2.5-72B-Instruct is a 72.7 billion parameter instruction-tuned causal language model developed by Qwen. It significantly improves upon its predecessor, Qwen2, with enhanced knowledge, coding, and mathematical capabilities, leveraging specialized expert models. This model excels at instruction following, long text generation (up to 8K tokens), structured data understanding, and JSON output, while supporting a 128K token context length and over 29 languages.

Loading preview...

Overview

Qwen2.5-72B-Instruct is the instruction-tuned variant of the latest Qwen2.5 series of large language models, developed by Qwen. This 72.7 billion parameter model builds upon the Qwen2 architecture, incorporating transformers with RoPE, SwiGLU, RMSNorm, and Attention QKV bias. It features 80 layers and a context length of up to 131,072 tokens, with a generation capacity of 8,192 tokens.

Key Capabilities and Improvements

  • Enhanced Knowledge & Specialized Skills: Significantly improved capabilities in coding and mathematics, attributed to the integration of specialized expert models.
  • Instruction Following & Text Generation: Demonstrates substantial improvements in adhering to instructions and generating long texts, particularly those exceeding 8,000 tokens.
  • Structured Data & Output: Better at understanding structured data like tables and generating structured outputs, especially JSON.
  • Robustness: More resilient to diverse system prompts, which enhances role-play implementations and chatbot condition-setting.
  • Long-Context Support: Supports a full context length of 131,072 tokens, utilizing techniques like YaRN for handling extensive inputs, though the default config.json is set for 32,768 tokens.
  • Multilingual Support: Provides support for over 29 languages, including major global languages like Chinese, English, French, Spanish, German, and Japanese.

Usage Notes

For processing long texts beyond 32,768 tokens, users can enable YaRN by modifying the rope_scaling configuration in config.json. It's noted that vLLM currently supports static YaRN, which might affect performance on shorter texts if enabled globally.