BytedTsinghua-SIA/Sequential-Qwen3-1.7B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 7, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Sequential-Qwen3-1.7B is a 1.7 billion parameter research checkpoint from the Qwen3 model family, developed by BytedTsinghua-SIA. It is specifically designed for reproducible research into post-training and reinforcement learning-style optimization for reasoning-oriented language models. This model is part of the Direct-OPD collection, focusing on studying optimization behaviors and serving as a comparison point for further research.

Loading preview...

Overview

Sequential-Qwen3-1.7B is a 1.7 billion parameter research checkpoint from the Qwen3 model family, developed by BytedTsinghua-SIA. It is part of the Direct-OPD collection and was initialized from a Qwen3 1.7B base, then further trained using a sequential Direct-OPD recipe.

Key Characteristics

  • Research-focused: Primarily intended for studying post-training and RL-style optimization, particularly for reasoning-oriented language models.
  • Provenance: A mirrored checkpoint from ModelScope, providing access to a specific training trajectory within the Direct-OPD framework.
  • Open License: Released under the Apache-2.0 license, facilitating research and development.

Intended Use Cases

This model is suitable for research activities such as:

  • Investigating post-training and RL-style optimization behaviors.
  • Comparing Direct-OPD checkpoints across different model families and parameter scales.
  • Performing offline evaluations on reasoning, instruction-following, and alignment benchmarks.
  • Serving as an initialization point or comparison baseline for new research.

Limitations

Users should be aware that this is a research checkpoint. It may generate incorrect or biased content, and full training details are not yet comprehensively documented. Performance can be sensitive to prompt formats and decoding settings, and research checkpoints might show regressions in general instruction-following while optimizing for specific objectives.