OzodovSanjar/Qwe-2.5-3b-coder

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 2, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

OzodovSanjar/Qwe-2.5-3b-coder is a 3.09 billion parameter causal language model from the Qwen2.5-Coder series, developed by Qwen. This model is specifically optimized for code generation, code reasoning, and code fixing, building upon the Qwen2.5 architecture. It features a full 32,768 token context length and maintains strong performance in mathematics and general competencies. This model is ideal for developers requiring robust coding capabilities in a compact package.

Loading preview...

Qwen2.5-Coder-3B Overview

OzodovSanjar/Qwe-2.5-3b-coder is a 3.09 billion parameter model from the Qwen2.5-Coder family, a series of code-specific large language models developed by Qwen. This model represents an advancement over CodeQwen1.5, with significant improvements in core coding tasks. It is pre-trained on an extensive 5.5 trillion tokens, including a substantial amount of source code, text-code grounding, and synthetic data, making it highly proficient in programming-related applications.

Key Capabilities and Features

  • Enhanced Code Performance: Demonstrates significant improvements in code generation, code reasoning, and code fixing.
  • Comprehensive Foundation: Designed to support real-world applications like Code Agents, while also maintaining strong mathematical and general competencies.
  • Architecture: Utilizes a transformer architecture with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings.
  • Context Length: Features a substantial 32,768 token context window, allowing for processing of larger codebases and complex prompts.
  • Model Size: This specific repository hosts the 3.09 billion parameter version, part of a family ranging from 0.5B to 32B parameters.

Recommended Use Cases

This model is primarily intended for tasks requiring strong coding abilities. While it is a base language model, it is well-suited for further fine-tuning (e.g., SFT, RLHF) or for fill-in-the-middle tasks. It is not recommended for direct conversational use without additional post-training.