ProCreations/grug-v2-9b
ProCreations/grug-v2-9b is a 9 billion parameter language model developed by ProCreations, specifically fine-tuned for improved internal reasoning and tool-use capabilities. This model excels at generating clean, concise internal thought processes while maintaining strong performance on coding and agentic tasks. It features a 'dialect fix' that refines its internal monologue, making it more efficient and less verbose, without compromising its external outputs or tool-calling accuracy. The model is optimized for use cases requiring robust code generation and complex tool-orchestration.
Loading preview...
ProCreations/grug-v2-9b: Enhanced Reasoning and Tool Use
ProCreations/grug-v2-9b is a 9 billion parameter model from ProCreations, distinguished by its unique "dialect fix" that refines its internal reasoning process. This update, applied to the private <think> target, significantly reduces verbosity and improves the clarity of the model's internal monologue without altering its external answers or tool calls. The model maintains strong performance across various coding and agentic benchmarks.
Key Capabilities & Improvements
- Refined Internal Reasoning: The dialect fix dramatically reduces function-word ratio and eliminates verbose internal traces like "User wants hello world Python," leading to more concise and efficient thought processes.
- Robust Coding Performance: The
grug-v2-9b correctedversion shows strong and improved performance on coding benchmarks, achieving 82.9% on HumanEval and 77.0% on MBPP. - Accurate Tool Use: Demonstrates 100% validity and strictness in tool calls across 'card' and 'broad' benchmarks, with high accuracy in selecting the 'right tool' (94.1% on 'broad').
- Stable Performance: The dialect fix was applied without degrading existing coding or tool-use capabilities, ensuring that improvements in internal style did not compromise external functionality.
Ideal Use Cases
- Code Generation: Excels in generating Python code, as evidenced by its benchmark scores.
- Agentic Workflows: Highly suitable for tasks requiring precise tool orchestration and complex multi-step reasoning, where clear internal thought processes are beneficial.
- Development Environments: Can be integrated into systems where a model's internal reasoning trace is monitored or used for debugging, benefiting from its concise and clean internal dialect.