Spreadsheet-RL/Spreadsheet-RL-8B
Spreadsheet-RL/Spreadsheet-RL-8B is an 8 billion parameter language model developed by Spreadsheet-RL, fine-tuned from Qwen/Qwen3-8B. This model is specifically post-trained using reinforcement learning (RL) within Spreadsheet Gym, a multi-turn Microsoft Excel environment. It excels at realistic spreadsheet tasks by leveraging spreadsheet-native tools, sandboxed code execution, and Excel-based recalculation rewards, making it ideal for agent-based automation in spreadsheet environments.
Loading preview...
Spreadsheet-RL-8B: An RL-Trained Agent for Excel Tasks
Spreadsheet-RL-8B is an 8 billion parameter model, originating from Qwen/Qwen3-8B, that has undergone specialized post-training using reinforcement learning (RL). Developed by Spreadsheet-RL, this model is designed to operate as an agent within complex Microsoft Excel environments, leveraging a unique setup called Spreadsheet Gym.
Key Capabilities & Features
- Reinforcement Learning (RL) Post-Training: Utilizes GRPO with outcome-based rewards, trained on 5,928 filtered ExcelForum tasks.
- Spreadsheet-Native Interaction: Integrated with Spreadsheet Gym, providing access to spreadsheet-native tools and sandboxed code execution.
- Enhanced Performance on SpreadsheetBench: Achieves a Pass@1 score of 22.3 on SpreadsheetBench, demonstrating a significant improvement over its base model and full-harness pre-RL results.
- Multi-Turn Excel Environment: Designed for multi-turn interactions within Microsoft Excel 365, enabling complex task automation.
When to Use This Model
This model is specifically intended for use with the Spreadsheet-RL agent harness and tool environment. It is particularly well-suited for:
- Developing and evaluating AI agents for realistic spreadsheet automation tasks.
- Research into reinforcement learning applications for complex, tool-augmented language models.
- Scenarios requiring precise interaction and problem-solving within Microsoft Excel, leveraging its unique training on Excel-specific tasks and tools.