eousphoros/qwen3.8-27b-yarn4-mixed-kl-128k-r1-iter137
The eousphoros/qwen3.8-27b-yarn4-mixed-kl-128k-r1-iter137 is an experimental 27 billion parameter merged derivative of the Qwen/Qwen3.8-27B model, enhanced with a static YaRN factor of 4.0 to extend its context window from 262,144 to 1,048,576 tokens. This model is designed to preserve the base Qwen3.8-27B's strong capabilities in coding, agentic tasks, and multimodal understanding while significantly increasing its long-context processing ability. It is particularly optimized for complex, long-horizon tasks requiring extensive context, such as advanced coding, professional work, and multimodal agentic intelligence.
Loading preview...
Overview
This model, eousphoros/qwen3.8-27b-yarn4-mixed-kl-128k-r1-iter137, is an experimental merged derivative of the Qwen/Qwen3.8-27B base model. It integrates a LoRA/direct-normalization calibration with a static YaRN factor of 4.0, extending the native context window from 262,144 tokens to 1,048,576 tokens. This modification aims to maintain the base model's performance while enabling significantly longer context processing.
Key Capabilities
- Extended Context Window: Processes up to 1,048,576 tokens, a four-fold increase over the base model's native 262,144 tokens, achieved via static YaRN scaling.
- Multimodal Understanding: Inherits native support for image and video understanding from the Qwen3.8-27B base.
- Enhanced Agentic Performance: Demonstrates strong capabilities in agentic tasks, including coding (e.g., SWE-bench Pro: 61.7, QwenSWEBench: 79.0), long-horizon office work (CoWorkBench: 70.7), and frontier agentic tasks (Agents' Last Exam Pass@1: 20.4).
- Flexible Thinking Control: Supports adjustable reasoning depth (
xhigh,medium,low) and preserves thinking context across messages.
What Makes This Different?
This model's primary differentiator is its significantly extended context length (1M tokens) through a static YaRN factor, making it suitable for applications requiring processing of very long documents, codebases, or video streams. While the base Qwen3.8-27B already offers a substantial context, this derivative pushes that boundary further, specifically targeting use cases where ultra-long context is critical. It's important to note that the long-context performance of this specific derivative has not yet undergone a full RULER evaluation, and the benchmark results provided are for the base Qwen3.8-27B model.
Should I Use This?
This model is ideal if your application requires processing extremely long sequences (up to 1 million tokens) and benefits from the Qwen3.8-27B's strong multimodal and agentic capabilities. Consider using this model for:
- Advanced code generation and software engineering tasks involving large repositories.
- Complex, multi-step agentic workflows that need to maintain extensive historical context.
- Multimodal applications analyzing hour-scale videos or large documents with embedded images.
However, as an experimental checkpoint with static YaRN, it's recommended to evaluate its performance for your specific workload before production deployment, especially for shorter texts where static YaRN might impact performance.