AllSpark-Research/CompoWorld
AllSpark-Research/CompoWorld is a 35.1 billion parameter Qwen3.6-35B-A3B agent model developed by AllSpark-Research, fine-tuned using the CompoWorld compositional environment scaling method. This model excels at complex agentic tasks by composing a finite library of reusable services, demonstrating a +9.17 point average improvement across eight agent benchmarks. It is specifically optimized for general agent capabilities, particularly in scenarios requiring information flow across multiple services, and features a 32768-token context length.
Loading preview...
CompoWorld: General Agent with Compositional Environment Scaling
CompoWorld is a 35.1 billion parameter Qwen3.6-35B-A3B agent model developed by AllSpark-Research, specifically trained to handle complex, multi-service agentic tasks. It leverages a novel compositional environment scaling method to expand the task space, enabling agents to connect information and actions across various services.
Key Capabilities & Training:
- Compositional Task Generation: Utilizes a random-walk procedure to connect services through dependency graphs, generating tasks that require information flow across multiple services.
- Verified Services: Employs coding agents to turn tool specifications into verified services with typed states and shared interfaces.
- Hybrid Training: Combines 3K supervised fine-tuning (SFT) trajectories and 1K reinforcement learning (RL) tasks, generated from 448 composed services exposing 10,130 tools.
- Completion-Focused Rubric Reward: Guides RL by emphasizing criteria with lower pass rates for full task completion.
- Base Model: Built upon Qwen/Qwen3.6-35B-A3B, a Mixture-of-Experts (MoE) model with a vision encoder and a 262,144-token context length.
Performance Highlights:
- Significant Improvement: Achieves an average improvement of +9.17 points over its backbone across eight challenging agent benchmarks.
- Frontier Model Performance: Surpasses frontier models like Claude Opus 4.6 on AutomationBench and leads all compared agent-specialized 35B-A3B models.
Good for:
- Developing general agents that need to interact with and integrate multiple services.
- Tasks requiring complex reasoning and information synthesis across different tools.
- Applications demanding robust performance on agent benchmarks like AutomationBench and SkillsBench.