unwrenchable/Qwen2.5-Coder-1.5B
unwrenchable/Qwen2.5-Coder-1.5B is a 1.54 billion parameter causal language model from the Qwen2.5-Coder series, developed by Qwen. This model is specifically pre-trained and optimized for code generation, code reasoning, and code fixing, building upon the strong Qwen2.5 foundation. It features a 32,768 token context length and utilizes a transformer architecture with RoPE, SwiGLU, and RMSNorm. The model is designed to enhance coding capabilities while maintaining strengths in mathematics and general competencies.
Loading preview...
Qwen2.5-Coder-1.5B Overview
Qwen2.5-Coder-1.5B is a 1.54 billion parameter model from the Qwen2.5-Coder series, a family of code-specific large language models developed by Qwen. This model is pre-trained and designed to significantly improve upon its predecessor, CodeQwen1.5, particularly in coding tasks. It is part of a larger series that includes models ranging from 0.5B to 32B parameters, with the 32B version reportedly matching GPT-4o in coding abilities.
Key Capabilities & Features
- Enhanced Code Performance: Offers significant improvements in code generation, code reasoning, and code fixing.
- Extensive Training Data: Pre-trained on 5.5 trillion tokens, including a substantial amount of source code, text-code grounding, and synthetic data.
- Robust Foundation: Aims to provide a comprehensive base for real-world applications like Code Agents, while also maintaining strong performance in mathematics and general language understanding.
- Technical Architecture: Built on a transformer architecture, incorporating features like RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings.
- Context Length: Supports a full context length of 32,768 tokens.
Intended Use
This 1.5B parameter base model is primarily intended for further fine-tuning (e.g., SFT, RLHF) or for specific tasks like fill-in-the-middle. It is not recommended for direct conversational use without additional post-training.