mercor/Qwen3.6-35B-A3B-Mercor

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 14, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

mercor/Qwen3.6-35B-A3B-Mercor is a 35.1 billion parameter language model, based on the Qwen3.6-35B-A3B architecture, developed by mercor. It has been specifically post-trained using Reinforcement Learning (RL) on APEX-Agents off-the-shelf datasets. This model is optimized for knowledge work agents, leveraging a 32768 token context length for complex tasks. Its primary strength lies in its specialized RL training for agentic applications.

Loading preview...

Model Overview

mercor/Qwen3.6-35B-A3B-Mercor is a 35.1 billion parameter language model derived from the Qwen3.6-35B-A3B architecture. Developed by mercor, this model distinguishes itself through its specialized post-training methodology.

Key Capabilities

  • Reinforcement Learning (RL) Post-Training: The model has undergone extensive RL post-training using APEX-Agents off-the-shelf datasets. This process enhances its ability to perform complex, agent-like tasks.
  • Optimized for Knowledge Work Agents: Its training regimen specifically targets applications requiring advanced knowledge processing and agentic behavior, making it suitable for automated knowledge work.
  • Context Length: Features a substantial context window of 32768 tokens, enabling it to handle and process large amounts of information for intricate tasks.

What Makes This Different?

Unlike many general-purpose LLMs, mercor/Qwen3.6-35B-A3B-Mercor's core differentiation lies in its targeted RL post-training on specific agent datasets. This specialization aims to produce a model highly effective in scenarios where autonomous decision-making and complex task execution are paramount, as detailed in mercor's blog post on "Training frontier knowledge work agents: a 397B RL training guide with SkyRL".

Should You Use This?

This model is particularly well-suited for use cases involving:

  • Developing and deploying knowledge work agents.
  • Applications requiring advanced reasoning and autonomous task execution.
  • Scenarios where RL-tuned performance on agentic datasets is a critical requirement.

If your application demands a model specifically optimized for agentic capabilities through advanced RL training, this model offers a specialized solution.