extreme1228/ScaleCUA-qwen3.5-osworld-sft

VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 9, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

The extreme1228/ScaleCUA-qwen3.5-osworld-sft is a 9 billion parameter Qwen3.5-based computer-use agent, developed by ScaleCUA, specifically fine-tuned as an SFT start-point checkpoint for OSWorld tasks. This model is designed to initialize online Reinforcement Learning for agents that interact with computer interfaces, leveraging techniques like VeriGen for verifiable GUI-task synthesis. It excels in enabling agents to perform complex computer-use tasks, making it suitable for research and development in automated UI interaction and agent training.

Loading preview...

ScaleCUA-qwen3.5-osworld-sft: A Computer-Use Agent for OSWorld

This model, developed by ScaleCUA, is a 9 billion parameter Qwen3.5-based computer-use agent specifically designed as an SFT (Supervised Fine-Tuning) start-point checkpoint for the OSWorld environment. It serves as the initial state for subsequent online Reinforcement Learning (RL) to create more advanced agents capable of interacting with computer interfaces.

Key Capabilities & Features

  • Computer-Use Agent: Optimized for tasks involving interaction with graphical user interfaces (GUIs).
  • OSWorld Integration: Specifically trained and intended for use within the OSWorld evaluation framework.
  • SFT Start-Point: Provides a strong foundation for further online RL training, as detailed in the associated research paper.
  • Advanced Training Techniques: Leverages methodologies like VeriGen (verifiable GUI-task synthesis), Frontier Sampling, and Visual Context Segmentation to scale online RL for computer-use agents.
  • Research-Oriented: Released for research purposes, inheriting the license of its base model, Qwen3.5-9B.

Good For

  • Developing and evaluating computer-use agents: Ideal for researchers and developers working on automated UI interaction.
  • Initializing online RL training: Provides a robust SFT checkpoint to begin Reinforcement Learning for agents in environments like OSWorld.
  • Exploring advanced agent training: Useful for understanding and implementing techniques for scaling computer-use agents.

For detailed information on the underlying research and full results, refer to the SCALECUA paper and the project's GitHub repository.