divaspoudel/iol-div
The divaspoudel/iol-div model is an instruction-tuned 7.61 billion parameter causal language model from the Qwen2.5 series, developed by Qwen. It features a transformer architecture with a context length of 32768 tokens, extendable up to 128K tokens using YaRN. This model significantly improves upon Qwen2 in knowledge, coding, mathematics, instruction following, and structured data understanding, offering multilingual support for over 29 languages.
Loading preview...
Overview
This repository hosts the instruction-tuned 7.61 billion parameter Qwen2.5 model, part of the latest Qwen large language model series developed by Qwen. It builds upon the Qwen2 architecture, incorporating transformers with RoPE, SwiGLU, RMSNorm, and Attention QKV bias. The model supports a default context length of 32,768 tokens, with the capability to extend up to 128,000 tokens for long texts using the YaRN technique.
Key Capabilities & Improvements
- Enhanced Knowledge & Reasoning: Significantly improved capabilities in coding and mathematics, leveraging specialized expert models.
- Instruction Following: Demonstrates substantial improvements in adhering to instructions, generating long texts (over 8K tokens), and understanding structured data like tables.
- Structured Output Generation: Excels at producing structured outputs, particularly JSON.
- Robustness: More resilient to diverse system prompts, enhancing role-play and condition-setting for chatbots.
- Multilingual Support: Offers comprehensive support for over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, and Vietnamese.
When to Use This Model
This model is particularly well-suited for applications requiring:
- Advanced coding and mathematical problem-solving.
- Precise instruction following and structured data processing.
- Generation of long, coherent texts and structured outputs (e.g., JSON).
- Multilingual interactions across a broad range of languages.
- Scenarios benefiting from extended context understanding, especially with YaRN for very long inputs.