SeanWang0027/qwen3-1.7b-sciworld-sft-gpt54mini-ep2

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 19, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The SeanWang0027/qwen3-1.7b-sciworld-sft-gpt54mini-ep2 is a 2 billion parameter Qwen3-based causal language model, fine-tuned using Supervised Fine-Tuning (SFT) on ReAct trajectories generated by gpt-5.4-mini. This model is specifically optimized for interactive problem-solving within the ScienceWorld environment, demonstrating a 13.25% success rate on ScienceWorld tasks. It is designed for agentic applications requiring sequential decision-making and action generation in simulated scientific contexts.

Loading preview...

Model Overview

This model, qwen3-1.7b-sciworld-sft-gpt54mini-ep2, is an intermediate checkpoint from a 3-epoch Supervised Fine-Tuning (SFT) run of the Qwen3-1.7B base model. It has been fine-tuned on 2059 full-episode ReAct trajectories, generated by gpt-5.4-mini (OpenAI API) within the ScienceWorld environment. The training data includes both successful and failed episodes, with 46035 supervised teacher turns.

Key Capabilities

  • ScienceWorld Task Performance: Achieves a 13.25% success rate on ScienceWorld test variations, significantly outperforming the base Qwen3-1.7B model (0.12%).
  • ReAct Trajectory Learning: Trained to follow Thought:\n...\n\nAction:\n<one command> patterns, enabling sequential reasoning and action generation.
  • Agentic Behavior: Optimized for agent-like interactions in simulated environments, specifically for tasks requiring open/close OBJ type actions.

Training Details

  • Data Source: 2059 full-episode ReAct trajectories from gpt-5.4-mini on ScienceWorld training tasks.
  • Methodology: Supervised Fine-Tuning (SFT) over 3 epochs, with this checkpoint representing the end of the second epoch (step 128).
  • Prompt Format: Utilizes a ReAct prompt format, with agent-gym's ScienceWorld instructions as user turns and model replies structured with Thought: and Action: components.

When to Use This Model

This model is particularly suited for research and development in:

  • Agentic AI: Exploring and building agents capable of interacting with simulated environments.
  • ScienceWorld Benchmarking: As a baseline or component for tasks within the ScienceWorld environment.
  • ReAct Pattern Generation: Applications requiring models to generate structured thoughts and actions in a sequential manner.