LARK-Lab/EnvFactory-8B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 19, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

EnvFactory-8B by LARK-Lab is an 8 billion parameter model, fine-tuned from Qwen3-8B, specifically designed for advanced tool-use capabilities. It leverages a novel EnvFactory framework for automated synthesis of executable tool environments and robust reinforcement learning. This model excels at enabling LLMs to perform complex, multi-turn tasks by interacting with real-world APIs, making it ideal for developing sophisticated agentic applications.

Loading preview...

EnvFactory-8B: Scaling Tool-Use Agents

EnvFactory-8B, developed by LARK-Lab, is an 8 billion parameter model built upon Qwen3-8B, specifically engineered to enhance Large Language Models (LLMs) with robust tool-use capabilities through Agentic Reinforcement Learning (Agentic RL). The core innovation is the EnvFactory framework, which automates the exploration, verification, and deployment of stateful, executable tool environments from authentic resources.

Key Capabilities

  • Automated Environment Synthesis: Discovers, validates, and deploys MCP-based tool environments from real-world APIs, significantly reducing manual effort.
  • Topology-Aware Trajectory Sampling: Generates natural, multi-turn tool-use trajectories that capture implicit human reasoning, crucial for effective agent training.
  • Robust RL Training: Utilizes verified environments and calibrated refinement for stable and effective reinforcement learning, leading to superior performance.
  • Scalable Architecture: Achieves improved tool-use performance with a significantly smaller set of environments (85 environments across 7 domains) compared to traditional methods.

Performance Highlights

EnvFactory-8B demonstrates notable improvements over its base model, Qwen3-8B, across various tool-use benchmarks. For instance, it shows an increase in BFCL Multi Turn scores from 41.25 to 49.00 and a substantial rise in MCP-Atlas Pass Rate from 5.15 to 13.75, indicating enhanced proficiency in complex tool interaction scenarios. The model was trained using a combination of Supervised Fine-Tuning (SFT) on 53.4k filtered trajectories and Reinforcement Learning (RL) on 3.09k trajectories, leveraging DeepSpeed ZeRO-3 and a forked VeRL framework.

Good For

  • Developing advanced AI agents that require interaction with external tools and APIs.
  • Applications demanding robust, multi-turn tool-use capabilities.
  • Research and development in agentic AI and reinforcement learning for LLMs.