OrionLLM/GRM-2.6-Plus-0628

Hugging Face
VISIONConcurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 7, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

OrionLLM/GRM-2.6-Plus-0628 is a 27-billion-parameter reasoning model built on the Qwen3.6 architecture, developed by OrionLLM. This model is optimized for difficult, high-complexity tasks, particularly excelling in long-horizon agentic workflows and structured reasoning. It delivers strong performance for its size, making it suitable for advanced problem-solving, coding, and local agentic applications.

Loading preview...

Overview

GRM-2.6-Plus-0628 is a 27-billion-parameter reasoning model from OrionLLM, built upon the Qwen3.6 architecture. It represents an update to the GRM-2.6-Plus model, specifically enhanced for difficult, high-complexity tasks and long-horizon agentic workflows. The model emphasizes structured reasoning to provide accurate, coherent, and reliable responses, aiming to compete with frontier models in its capability while remaining efficient for practical deployment.

Key Capabilities

  • Elite-Level Reasoning: Optimized for complex reasoning workloads with strong step-by-step problem-solving.
  • Improved Agentic Performance: Enhanced for multi-step agentic tasks, maintaining coherence over extended workflows.
  • High Performance for Size: Delivers excellent capability relative to its 27B parameter count, balancing intelligence with practical deployment.
  • Advanced Coding and Agentic Use: Well-suited for code generation, structured problem-solving, and local agentic applications.
  • Practical Deployment: Designed to be efficient and usable on capable consumer and workstation hardware.

Performance Highlights

GRM-2.6-Plus-0628 demonstrates strong performance across various benchmarks, often outperforming its predecessor GRM-2.6-Plus and other models in its class like Qwen3.6-27B and google/gemma-4-31B-it. Notable scores include:

  • MMLU-Pro: 88.1
  • MMLU-Redux: 96.4
  • C-Eval: 92.4
  • GPQA Diamond: 90.1
  • LiveCodeBench v6: 86.5
  • SWE-bench Verified: 79.7
  • Terminal-Bench 2.0: 62.6

These benchmarks highlight its practical intelligence, strong task understanding, and stable responses across knowledge, STEM, reasoning, coding, and general agent tasks.

When to Use This Model

This model is ideal for developers and researchers who require:

  • A powerful local AI model for complex reasoning and problem-solving.
  • Robust capabilities for long-horizon agentic tasks and multi-step workflows.
  • Strong performance in code generation and structured problem-solving.
  • A balance of high intelligence and practical deployment efficiency on consumer hardware.