SeanWang0027/qwen3-1.7b-sciworld-sft-gpt54mini-ep1

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 19, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The SeanWang0027/qwen3-1.7b-sciworld-sft-gpt54mini-ep1 is a 2 billion parameter Qwen3-based model, fine-tuned for interactive problem-solving within the ScienceWorld environment. This intermediate checkpoint, from the first epoch of supervised fine-tuning, is designed to process and generate ReAct trajectories for complex scientific tasks. It specializes in agent-based reasoning and action generation, leveraging a 32768 token context length for detailed interaction sequences.

Loading preview...

Model Overview

This model, qwen3-1.7b-sciworld-sft-gpt54mini-ep1, is an intermediate checkpoint from the first epoch of a supervised fine-tuning (SFT) process. It is based on the Qwen3-1.7B architecture and has been specifically trained for the ScienceWorld environment, focusing on generating ReAct trajectories.

Training Details

  • Data: The model was trained on 2059 full-episode ReAct trajectories generated by gpt-5.4-mini (OpenAI API) on ScienceWorld training tasks. This dataset includes 46035 supervised teacher turns, encompassing both successful and failed episodes, and rejected actions.
  • Methodology: Training involved one row per episode, with cross-entropy loss applied only to each teacher reply and its <|im_end|>. It was part of a 3-epoch run, utilizing a batch size of 32, AdamW optimizer with a learning rate of 1e-5, and a cosine learning rate schedule with 10% warmup over 192 steps. The model was trained in fp32 with a maximum length of 8192 tokens.

Current Status

This specific checkpoint represents step 64 of the training and is stored in bf16. It is important to note that this intermediate checkpoint has not been evaluated for performance. A separate, fully trained single-epoch run achieved a 9.12% success rate on the ScienceWorld test set.

Prompt Format

The model uses a ReAct prompt format, where AgentGym's ScienceWorld instructions are provided as a user turn, followed by a canned assistant acknowledgment, and then one user turn per observation. The model is expected to reply with Thought:\n...\n\nAction:\n<one command>. It is designed to be rendered with Qwen3's chat template with enable_thinking=False.