bytkim/Qwen3.8-27B-pi
bytkim/Qwen3.8-27B-pi is a 27 billion parameter language model, fine-tuned from Qwen3.8-27B, specifically optimized for coding tasks within the Pi agent harness. It excels at iterative coding workflows, including repository reading, file editing, tool execution, and feedback integration, with an emphasis on efficient problem-solving. The model features an extended context length of 32768 tokens and is designed to produce more completed work with less generated text.
Loading preview...
Model Overview
bytkim/Qwen3.8-27B-pi is a 27 billion parameter model, fine-tuned from Qwen3.8-27B, with a 32768 token context length. It is specifically engineered for coding work within the Pi agent harness, focusing on the iterative process of reading repositories, editing files, running tools, and responding to feedback. The model aims to achieve more completed work with less generated text, while maintaining the base model's interface.
Key Capabilities & Features
- Optimized for Pi Agent Harness: Fine-tuned to handle complete coding workflows, including adapting to existing environments and checking results against task requirements.
- Efficient Coding Workflows: Learns to turn plans into working, checked implementations through supervised fine-tuning on successful Pi sessions.
- Adjustable Reasoning Effort: Balances responsiveness with deeper problem-solving, allowing for focused edits or more involved debugging.
- Reinforcement Learning Refinement: Utilizes a custom reasoning-efficiency reward built on GRPO to encourage economical solutions for low- and medium-effort tasks, while prioritizing correctness for high-effort tasks.
- Practical Formats: Available in various deployment formats, including BF16, FP8, and GGUF quantizations, calibrated using complete Pi coding sessions.
Performance Highlights
Benchmarks against the base Qwen3.8-27B model show that Qwen3.8-27B-pi achieves a steadier rise in task completion across reasoning levels (low, medium, xhigh) on Terminal-Bench 2.1, often with fewer output tokens. For instance, its medium setting can match the base model's xhigh completion rate with approximately 41% fewer output tokens. On GPQA Diamond, Pi shows a smoother increase in attempt success with effort, achieving the highest xhigh score. For SciCode, Pi solves more subproblems at every reasoning level, scoring higher at xhigh with about 23% fewer output tokens.
Good for
- Developers and teams using or building AI agents for coding tasks.
- Applications requiring iterative code development, debugging, and tool interaction.
- Scenarios where efficient resource usage and concise, effective outputs are critical for coding solutions.