AllSpark-Research/CompoWorld

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 29, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

AllSpark-Research/CompoWorld is a 35.1 billion parameter Qwen3.6-35B-A3B agent model developed by AllSpark-Research, fine-tuned using the CompoWorld compositional environment scaling method. This model excels at complex agentic tasks by composing a finite library of reusable services, demonstrating a +9.17 point average improvement across eight agent benchmarks. It is specifically optimized for general agent capabilities, particularly in scenarios requiring information flow across multiple services, and features a 32768-token context length.

Loading preview...

CompoWorld: General Agent with Compositional Environment Scaling

CompoWorld is a 35.1 billion parameter Qwen3.6-35B-A3B agent model developed by AllSpark-Research, specifically trained to handle complex, multi-service agentic tasks. It leverages a novel compositional environment scaling method to expand the task space, enabling agents to connect information and actions across various services.

Key Capabilities & Training:

  • Compositional Task Generation: Utilizes a random-walk procedure to connect services through dependency graphs, generating tasks that require information flow across multiple services.
  • Verified Services: Employs coding agents to turn tool specifications into verified services with typed states and shared interfaces.
  • Hybrid Training: Combines 3K supervised fine-tuning (SFT) trajectories and 1K reinforcement learning (RL) tasks, generated from 448 composed services exposing 10,130 tools.
  • Completion-Focused Rubric Reward: Guides RL by emphasizing criteria with lower pass rates for full task completion.
  • Base Model: Built upon Qwen/Qwen3.6-35B-A3B, a Mixture-of-Experts (MoE) model with a vision encoder and a 262,144-token context length.

Performance Highlights:

  • Significant Improvement: Achieves an average improvement of +9.17 points over its backbone across eight challenging agent benchmarks.
  • Frontier Model Performance: Surpasses frontier models like Claude Opus 4.6 on AutomationBench and leads all compared agent-specialized 35B-A3B models.

Good for:

  • Developing general agents that need to interact with and integrate multiple services.
  • Tasks requiring complex reasoning and information synthesis across different tools.
  • Applications demanding robust performance on agent benchmarks like AutomationBench and SkillsBench.