fanie123/Qwen2.5-Coder-14B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:14.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 23, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The fanie123/Qwen2.5-Coder-14B is a 14.7 billion parameter causal language model from the Qwen2.5-Coder series, developed by Qwen. This model is specifically designed for advanced code generation, code reasoning, and code fixing, building upon the Qwen2.5 architecture. It features a transformer architecture with RoPE, SwiGLU, and RMSNorm, and supports a full context length of 131,072 tokens. Optimized for coding tasks, it also maintains strong performance in mathematics and general competencies, making it suitable for complex real-world applications like Code Agents.

Loading preview...

Qwen2.5-Coder-14B Overview

Qwen2.5-Coder-14B is a 14.7 billion parameter causal language model, part of the Qwen2.5-Coder series developed by Qwen. This series represents an evolution from CodeQwen1.5, focusing on significant enhancements in coding capabilities. The model's training dataset was scaled up to 5.5 trillion tokens, incorporating source code, text-code grounding, and synthetic data, to improve code generation, code reasoning, and code fixing.

Key Capabilities

  • Enhanced Coding Performance: Demonstrates substantial improvements in various code-related tasks, with the 32B variant reportedly matching GPT-4o's coding abilities.
  • Broad Application Foundation: Designed to support real-world applications such as Code Agents, while retaining strong performance in mathematics and general competencies.
  • Extended Context Length: Supports a full context length of 131,072 tokens, with the ability to handle even longer texts up to 128K tokens using techniques like YaRN for extrapolation.
  • Robust Architecture: Built on a transformer architecture featuring RoPE, SwiGLU, RMSNorm, and Attention QKV bias.

Good For

  • Developers requiring a powerful base model for code generation, reasoning, and fixing.
  • Building Code Agents or other applications that demand strong coding and general AI capabilities.
  • Tasks involving long codebases or extensive textual inputs due to its large context window.

It is important to note that this is a base language model, and for conversational use cases, post-training methods like SFT or RLHF are recommended.