allenai/qwen35-9b-openthoughts

VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 19, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The allenai/qwen35-9b-openthoughts model is a 9 billion parameter language model developed by Ai2, fine-tuned from Qwen 3.5 9B using DPPO. It is specifically optimized as a terminal agent, trained on the OpenThoughts-Agent dataset. This model excels at terminal-based tasks, demonstrating strong performance on the TBLite benchmark with a score of 53.0 ± 0.7.

Loading preview...

Overview

This model, allenai/qwen35-9b-openthoughts, is a 9 billion parameter language model developed by Ai2. It is fine-tuned from the Qwen 3.5 9B base model using Deep Proximal Policy Optimization (DPPO) and is specifically designed to function as a terminal agent. The training utilized the OpenThoughts-Agent dataset as an ablation study within a larger collection of terminal agents.

Key Capabilities & Performance

  • Terminal Agent Specialization: Optimized for interacting with and performing tasks within a terminal environment.
  • DPPO Fine-tuning: Leverages DPPO for enhanced agentic capabilities.
  • Benchmark Performance: Achieves a score of 53.0 ± 0.7 on the TBLite benchmark, indicating strong proficiency in terminal-based tasks. This places it competitively among other Qwen 3.5 9B variants fine-tuned for similar purposes.
  • Context Length: Supports a maximum prompt length of 2048 tokens and a maximum per-turn length of 16384 tokens, with an overall maximum of 65536 tokens.

Training Details

The model was trained for 200 steps (out of a total 500) using specific hyperparameters, including a learning rate of 1e-6 and a constant LR scheduler. More details on the training methodology can be found in the associated Tmax paper and codebase.

License

This model is released under the Apache 2.0 License, intended for research and educational use in line with Ai2's Responsible Use Guidelines.