cocoa-org/Mocha-Coder-32B

TEXT GENERATIONPricing:Input $2.72 / Output $4.8Concurrent Unit Cost:2Model Size:32.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 13, 2026License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Mocha-Coder-32B is a 32.8 billion parameter coding agent developed by cocoa-org, built upon Qwen2.5-Coder-32B-Instruct. It is trained through distillation on a 300K+ trajectory mixture using the NanoRollout infrastructure, without reinforcement learning. This model excels in agentic software engineering (SWE) tasks, achieving state-of-the-art performance among open-data models at its scale and competing with much larger open-source models on benchmarks like SWE-Bench Verified and SWE-Bench Pro.

Loading preview...

Mocha-Coder-32B: A Strong Open-Data Coding Agent

Mocha-Coder-32B is a 32.8 billion parameter coding agent developed by cocoa-org, built on Qwen2.5-Coder-32B-Instruct. It is trained entirely through distillation on a 300K+ trajectory mixture, sampled using the lightweight agent-rollout infrastructure, NanoRollout, without reinforcement learning. This model leverages training signals from frontier open-source teacher models (Qwen3-Coder-480B-A35B, Kimi-K2.5, Qwen3-Coder-Next, DeepSeek-V3.2) across multiple agent harnesses (OpenHands, mini-swe-agent, Terminus-2 JSON) on SWE-Rebench, SWE-Smith, and SETA benchmarks.

Key Capabilities

  • Strong Agentic SWE Performance: Achieves 62.6 Pass@1 on SWE-Bench Verified, 35.3 on SWE-Bench Pro, and 23.6 on Terminal-Bench 2.0, making it competitive with larger models like Qwen3-Coder-480B-A35B-Instruct.
  • Multi-Harness Training: Trajectories cover OpenHands, mini-swe-agent, and Terminus-2 JSON, which helps mitigate harness-specific overfitting.
  • Open Data: Distilled from a fully released 300K+ trajectory mixture (ZeonLap/Mocha-trajectories), promoting transparency and reproducibility.

Good For

  • Agentic Software Engineering Tasks: Ideal for use cases requiring automated code generation, debugging, and problem-solving within agentic frameworks.
  • Integration with Coding Agent Harnesses: Designed to be paired with harnesses like mini-swe-agent, OpenHands, or Terminus-2 JSON for optimal performance.
  • Research and Development: Suitable for researchers and developers exploring distillation techniques and agentic model performance on open-source benchmarks.