Darth-Coder/my-model2.5-14b-it
Darth-Coder/my-model2.5-14b-it is a 14.7 billion parameter instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen. It features significant improvements in coding, mathematics, instruction following, and long text generation up to 8K tokens, with a full context length of 131,072 tokens. This model excels at understanding structured data like tables and generating structured outputs, including JSON, while supporting over 29 languages.
Loading preview...
Qwen2.5-14B-Instruct Overview
Darth-Coder/my-model2.5-14b-it is an instruction-tuned variant of the Qwen2.5 series, a powerful family of large language models developed by Qwen. This specific model features 14.7 billion parameters and is built upon a transformer architecture incorporating RoPE, SwiGLU, RMSNorm, and Attention QKV bias. A key highlight is its extensive context length of 131,072 tokens, with the ability to generate responses up to 8,192 tokens.
Key Capabilities
- Enhanced Coding and Mathematics: Significantly improved performance in these domains due to specialized expert models.
- Superior Instruction Following: More robust instruction adherence and better handling of diverse system prompts for role-play and condition-setting.
- Advanced Text Generation: Excels at generating long texts (over 8K tokens) and understanding/generating structured data, including JSON.
- Multilingual Support: Supports over 29 languages, such as Chinese, English, French, Spanish, German, and Japanese.
- Long-Context Processing: Utilizes YaRN for handling inputs exceeding 32,768 tokens, with a default configuration for 32,768 tokens.
Good for
- Applications requiring strong coding and mathematical reasoning.
- Chatbots and agents needing robust instruction following and role-play capabilities.
- Tasks involving long document summarization or generation.
- Generating structured outputs like JSON from natural language prompts.
- Multilingual applications across a broad range of languages.