unsloth/Qwen2.5-3B

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 18, 2024License:otherArchitecture:Transformer0.0K Featherless Exclusive Warm

unsloth/Qwen2.5-3B is a 3.09 billion parameter causal language model from the Qwen2.5 series, developed by Qwen. This base model features a 32,768-token context length and is designed for pretraining, offering significantly improved capabilities in coding, mathematics, instruction following, and generating structured outputs like JSON. It supports over 29 languages and is intended for further fine-tuning rather than direct conversational use.

Loading preview...

Qwen2.5-3B: A Foundation Model for Advanced Fine-tuning

unsloth/Qwen2.5-3B is a 3.09 billion parameter base causal language model from the Qwen2.5 series, developed by Qwen. It is built on a transformer architecture with RoPE, SwiGLU, RMSNorm, and a 32,768-token context length. This model is specifically designed for pretraining and serves as a robust foundation for subsequent fine-tuning tasks such as SFT or RLHF.

Key Capabilities & Improvements

Qwen2.5 models, including this 3B variant, offer significant enhancements over previous Qwen2 versions:

  • Enhanced Knowledge & Specialized Skills: Greatly improved capabilities in coding and mathematics, leveraging specialized expert models.
  • Instruction Following: Substantial improvements in adhering to instructions and generating long texts (up to 8K tokens).
  • Structured Data & Output: Better understanding of structured data (e.g., tables) and improved generation of structured outputs, particularly JSON.
  • Multilingual Support: Supports over 29 languages, including Chinese, English, French, Spanish, German, and Japanese.
  • Robustness: More resilient to diverse system prompts, enhancing role-play and condition-setting for chatbots.

Intended Use

This 3B model is a base language model and is not recommended for direct conversational use. Its primary purpose is to be fine-tuned for specific applications. Developers can leverage its advanced pretraining for tasks requiring strong coding, mathematical reasoning, or structured output generation after applying post-training techniques.