konizquants/OpenThinker-Agent-v1-SFT
The OpenThinker-Agent-v1-SFT model by OpenThoughts is an 8 billion parameter language model, post-trained from Qwen3-8B. It is specifically fine-tuned for agentic tasks, excelling in environments like Terminal-Bench 2.0 and SWE-Bench. This model represents the supervised fine-tuning (SFT) stage of the OpenThinker-Agent-v1 development, focusing on curating high-quality datasets for agent training.
Loading preview...
OpenThinker-Agent-v1-SFT: Agentic Model for Complex Tasks
OpenThinker-Agent-v1-SFT is an 8 billion parameter model developed by OpenThoughts, derived from the Qwen3-8B architecture. This model is the result of the supervised fine-tuning (SFT) stage within the broader OpenThinker-Agent-v1 project, which aims to create robust models for agentic tasks.
Key Capabilities & Training
- Agentic Task Specialization: Fine-tuned for performance on complex agentic benchmarks such as Terminal-Bench 2.0 and SWE-Bench.
- Supervised Fine-Tuning (SFT): Trained on the OpenThoughts-Agent-v1-SFT dataset, which comprises approximately 15,200 traces. This dataset includes tasks like
nl2bash(shell command formatting) andInferredBugs(C# and Java bug-fixing tasks). - Data Curation: The OpenThoughts project emphasizes curating high-quality datasets for agent training, utilizing a three-stage filtration pipeline to ensure data stability and quality.
Performance Context
While this specific model is the SFT stage, the full OpenThinker-Agent-v1 (which includes an additional Reinforcement Learning stage) demonstrates significant improvements over its base model, Qwen3-8B, on agent benchmarks. For instance, the full RL-trained model achieves 15.7% on SWE-Bench Verified compared to 0.7% for Qwen3-8B.
Use Cases
This model is particularly well-suited for developers and researchers focused on:
- Developing AI agents: As a foundational SFT model for agentic workflows.
- Automated code generation and debugging: Especially for tasks involving shell commands and bug resolution in programming languages.
- Research in agentic AI: Providing a strong base for further experimentation and reinforcement learning.