OrionLLM/GRM-Coder-14b

TEXT GENERATIONPricing:Input $0.48 / Output $0.96Concurrent Unit Cost:1Model Size:14BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 18, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

OrionLLM/GRM-Coder-14b is a 14 billion parameter coding model based on Qwen3-14B, specifically designed and optimized for competitive programming tasks. It achieves a Pass@1 accuracy of 67.87% on LiveCodeBench v6, representing a 7.08% improvement over its Qwen3-14B base. The model was trained on 24,000 verifiable coding problems, making it highly proficient in generating solutions for complex programming challenges. With a context length of 32768 tokens, it can handle extensive code inputs and problem descriptions.

Loading preview...

OrionLLM/GRM-Coder-14b: A Specialized Competitive Programming Model

OrionLLM/GRM-Coder-14b is a powerful 14 billion parameter coding model built upon the Qwen3-14B architecture. Its primary focus is competitive programming, where it demonstrates significant performance gains over its base model.

Key Capabilities and Performance

  • Enhanced Competitive Programming Performance: The model achieves a Pass@1 accuracy of 67.87% on LiveCodeBench v6 (evaluated between 08/01/2024 - 05/01/2025). This marks a substantial 7.08% improvement compared to the baseline Qwen3-14B's Pass@1 accuracy of 60.79%.
  • Extensive Training Data: GRM-Coder-14b was trained on 24,000 verifiable coding problems over a four-day period, specifically targeting the nuances and complexities of competitive programming challenges.
  • Large Context Window: With a context length of 32768 tokens, it can process and generate solutions for problems requiring extensive code or detailed problem descriptions.

When to Use This Model

  • Competitive Programming: Ideal for tasks involving generating solutions for competitive programming problems.
  • Code Generation: Suitable for general code generation where high accuracy and problem-solving capabilities are paramount.
  • Benchmarking: Useful for researchers and developers looking for a strong baseline or specialized model in competitive programming benchmarks.