beyoru/Kiwen1.1-27B
Kiwen1.1-27B by beyoru is a 27 billion parameter language model, fine-tuned from Qwen3.8-27B with a 32768 token context length. It is specifically optimized for enhanced reasoning, coding, tool use, and instruction following, leveraging long chain-of-thought reasoning traces from Kimi K3. This model demonstrates significant improvements in mathematical reasoning, achieving 96.4% on GSM8K strict, and shows strong performance in instruction adherence and multi-turn tool calling.
Loading preview...
Overview
Kiwen1.1-27B is a 27 billion parameter model developed by beyoru, fine-tuned from Qwen3.8-27B. The primary goal of this model is to improve reasoning, instruction following, and reliability within a 27B parameter footprint, rather than simply increasing model size. It incorporates long chain-of-thought reasoning traces from Kimi K3, with a particular focus on coding and tool use.
Key Capabilities & Performance
- Enhanced Reasoning: Achieves 96.4% on GSM8K strict and 96.7% on GSM8K flexible, a substantial improvement over its base model.
- Instruction Following: Demonstrates strong performance in IFEval benchmarks, with 87.5% on IFEval instruction strict and 89.7% on IFEval instruction loose.
- Coding Focus: Training has been specifically directed towards coding tasks, improving its utility in programming contexts.
- Tool Use: Excels in multi-turn tool calling, showing a significant delta of +15.1 in internal benchmarks.
- Flexible Inference: Supports an
enable_thinkingflag during inference, allowing users to switch between detailed reasoning and direct extraction/classification modes.
Use Cases
Kiwen1.1-27B is particularly well-suited for applications requiring robust mathematical reasoning, precise instruction adherence, and effective code generation or tool interaction. Its optimized reasoning capabilities make it a strong candidate for complex problem-solving tasks where a smaller, more efficient model is desired.