Klingspor/Qwen3-4B-SFT

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 13, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Klingspor/Qwen3-4B-SFT is a 4 billion parameter supervised fine-tuned (SFT) version of Qwen3-4B, specifically designed for the '20 Questions' task. This model acts as a questioner, asking yes-or-no questions to deduce a secret word, and is optimized for multi-turn interactive language agent research. It serves as an initialization checkpoint for reinforcement learning models like StarPO and CIA, demonstrating specialized interactive reasoning capabilities within a 32768 token context.

Loading preview...

Overview of Klingspor/Qwen3-4B-SFT

This model is a supervised fine-tuned (SFT) version of Qwen3-4B, developed as part of the research presented in the paper "Intrinsic Credit Assignment for Long Horizon Interaction." Its primary function is to act as a Questioner in the game of 20 Questions, where it asks up to 20 yes-or-no questions to identify a secret common English noun.

Key Capabilities & Features

  • Specialized for 20 Questions: Designed to strategically ask deductive questions in a multi-turn interactive setting.
  • Reinforcement Learning Initialization: Serves as a crucial starting checkpoint for further reinforcement learning (RL) models, such as StarPO and CIA, focusing on long-horizon interaction.
  • Interactive Agent Research: Intended for research into multi-turn interactive language agents, providing a foundation for complex dialogue systems.

Training Details

The model was fine-tuned using a supervised approach on successful, filtered single-turn trajectories. The training data comprised 341 words from the COCA word list, ensuring no overlap with RL training or test sets. A Qwen3-14B model acted as the judge/oracle during this process.

Intended Use Cases

  • Playing 20 Questions: Directly usable as an agent to play the 20 Questions game.
  • RL Training Base: Ideal as an initial checkpoint for developing and training advanced RL-based interactive agents.
  • Research: Supports research into interactive language models and intrinsic credit assignment.

For more technical details, refer to the Intrinsic Credit Assignment for Long Horizon Interaction paper and the associated GitHub repository.