unwrenchable/Qwen2.5-Coder-1.5B

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

unwrenchable/Qwen2.5-Coder-1.5B is a 1.54 billion parameter causal language model from the Qwen2.5-Coder series, developed by Qwen. This model is specifically pre-trained and optimized for code generation, code reasoning, and code fixing, building upon the strong Qwen2.5 foundation. It features a 32,768 token context length and utilizes a transformer architecture with RoPE, SwiGLU, and RMSNorm. The model is designed to enhance coding capabilities while maintaining strengths in mathematics and general competencies.

Loading preview...

Qwen2.5-Coder-1.5B Overview

Qwen2.5-Coder-1.5B is a 1.54 billion parameter model from the Qwen2.5-Coder series, a family of code-specific large language models developed by Qwen. This model is pre-trained and designed to significantly improve upon its predecessor, CodeQwen1.5, particularly in coding tasks. It is part of a larger series that includes models ranging from 0.5B to 32B parameters, with the 32B version reportedly matching GPT-4o in coding abilities.

Key Capabilities & Features

  • Enhanced Code Performance: Offers significant improvements in code generation, code reasoning, and code fixing.
  • Extensive Training Data: Pre-trained on 5.5 trillion tokens, including a substantial amount of source code, text-code grounding, and synthetic data.
  • Robust Foundation: Aims to provide a comprehensive base for real-world applications like Code Agents, while also maintaining strong performance in mathematics and general language understanding.
  • Technical Architecture: Built on a transformer architecture, incorporating features like RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings.
  • Context Length: Supports a full context length of 32,768 tokens.

Intended Use

This 1.5B parameter base model is primarily intended for further fine-tuning (e.g., SFT, RLHF) or for specific tasks like fill-in-the-middle. It is not recommended for direct conversational use without additional post-training.