SeanWang0027/qwen3-1.7b-sciworld-sft-gpt54mini-ep2
The SeanWang0027/qwen3-1.7b-sciworld-sft-gpt54mini-ep2 is a 2 billion parameter Qwen3-based causal language model, fine-tuned using Supervised Fine-Tuning (SFT) on ReAct trajectories generated by gpt-5.4-mini. This model is specifically optimized for interactive problem-solving within the ScienceWorld environment, demonstrating a 13.25% success rate on ScienceWorld tasks. It is designed for agentic applications requiring sequential decision-making and action generation in simulated scientific contexts.
Loading preview...
Model Overview
This model, qwen3-1.7b-sciworld-sft-gpt54mini-ep2, is an intermediate checkpoint from a 3-epoch Supervised Fine-Tuning (SFT) run of the Qwen3-1.7B base model. It has been fine-tuned on 2059 full-episode ReAct trajectories, generated by gpt-5.4-mini (OpenAI API) within the ScienceWorld environment. The training data includes both successful and failed episodes, with 46035 supervised teacher turns.
Key Capabilities
- ScienceWorld Task Performance: Achieves a 13.25% success rate on ScienceWorld test variations, significantly outperforming the base Qwen3-1.7B model (0.12%).
- ReAct Trajectory Learning: Trained to follow
Thought:\n...\n\nAction:\n<one command>patterns, enabling sequential reasoning and action generation. - Agentic Behavior: Optimized for agent-like interactions in simulated environments, specifically for tasks requiring
open/close OBJtype actions.
Training Details
- Data Source: 2059 full-episode ReAct trajectories from
gpt-5.4-minion ScienceWorld training tasks. - Methodology: Supervised Fine-Tuning (SFT) over 3 epochs, with this checkpoint representing the end of the second epoch (step 128).
- Prompt Format: Utilizes a ReAct prompt format, with agent-gym's ScienceWorld instructions as user turns and model replies structured with
Thought:andAction:components.
When to Use This Model
This model is particularly suited for research and development in:
- Agentic AI: Exploring and building agents capable of interacting with simulated environments.
- ScienceWorld Benchmarking: As a baseline or component for tasks within the ScienceWorld environment.
- ReAct Pattern Generation: Applications requiring models to generate structured thoughts and actions in a sequential manner.