TreezzZ/opd-teacher-alfworld-7b
TreezzZ/opd-teacher-alfworld-7b is an instruction-tuned 7.61 billion parameter causal language model from the Qwen2.5 series, developed by Qwen. It features a transformer architecture with RoPE, SwiGLU, and RMSNorm, supporting a context length of up to 131,072 tokens. This model significantly enhances capabilities in coding, mathematics, instruction following, and generating structured outputs like JSON, making it suitable for complex reasoning and multilingual applications across 29 languages.
Loading preview...
Qwen2.5-7B-Instruct: Enhanced Multilingual LLM
This model, TreezzZ/opd-teacher-alfworld-7b, is an instruction-tuned variant of the Qwen2.5 series, developed by Qwen. It builds upon the Qwen2 architecture, offering substantial improvements across several key areas. With 7.61 billion parameters, it is designed for robust performance in diverse NLP tasks.
Key Capabilities and Improvements
- Enhanced Knowledge & Reasoning: Significantly improved capabilities in coding and mathematics, leveraging specialized expert models.
- Instruction Following: Better adherence to instructions, more resilient to varied system prompts, and improved role-play implementation.
- Long Text Generation & Understanding: Excels at generating long texts (up to 8K tokens) and understanding structured data like tables, with full context support up to 131,072 tokens.
- Structured Output Generation: Highly effective at producing structured outputs, particularly JSON.
- Multilingual Support: Provides robust support for over 29 languages, including major global languages like Chinese, English, French, Spanish, German, and Japanese.
- Architecture: Utilizes a transformer architecture with RoPE, SwiGLU, RMSNorm, and Attention QKV bias.
Long Context Handling
The model's config.json is set for a 32,768 token context length by default. For inputs exceeding this, it can be configured with YaRN (Yet another RoPE-scaling for Long-context) to extend context handling up to 131,072 tokens. This allows for processing very extensive documents, though static YaRN in deployments like vLLM might impact performance on shorter texts if always enabled.
Ideal Use Cases
This model is well-suited for applications requiring advanced coding assistance, mathematical problem-solving, complex instruction following, generation of structured data, and processing or generating long-form multilingual content.