OrionLLM/GRM-3.2-Cliff
OrionLLM/GRM-3.2-Cliff is a 9 billion parameter model built on the Ornith-1.0-9B architecture, specifically optimized for long-horizon agentic tasks and extremely difficult reasoning problems. It excels at complex coding challenges, advanced mathematics, and rigorous logical reasoning, designed to run efficiently in local, lower GPU environments. This model maintains coherence and planning quality across multi-step workflows, making it suitable for resource-constrained hardware. It represents a significant advancement in multi-step planning and self-correction performance for local execution.
Loading preview...
OrionLLM/GRM-3.2-Cliff: Agentic Reasoning for Local Environments
GRM-3.2-Cliff is a 9 billion parameter model developed by OrionLLM, built on the Ornith-1.0-9B architecture. It is an intermediate model within the GRM-3.2 family, specifically engineered for long-horizon agentic tasks and extremely difficult reasoning problems in local environments. This model significantly improves long-horizon task capability over its predecessor, GRM-2.5-Plus, by addressing common failures like contextual drift and multi-step degradation.
Key Capabilities
- Long-Horizon Agentic Mastery: Optimized for maintaining coherence, planning quality, and task fidelity across extended, multi-step agentic workflows.
- Local Workflow Efficiency: Designed to run smoothly in lower GPU environments while delivering high-tier reasoning performance.
- Elite Reasoning: Strong performance on challenging coding, advanced mathematics, and logical reasoning tasks through structured, step-by-step problem-solving.
- Robust Coding Ability: Handles complex, multi-file coding tasks, debugging, refactoring, and long-running terminal sessions locally.
Performance Highlights
GRM-3.2-Cliff demonstrates strong performance in agentic benchmarks. It achieves 70.3 on SWE-bench Verified for agentic coding and 43.4 on SWE-bench Pro for real-world software engineering, showing an improvement over GRM-2.5-Plus (35.6). For agentic terminal coding, it scores 45.3 on Terminal-Bench 2.1, a substantial increase from GRM-2.5-Plus (27.8). It also shows a unique strength in repo-level code generation, scoring 28.5 on NL2Repo, a benchmark not reported for comparable models.