BytedTsinghua-SIA/QuestA-Qwen3-1.7B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

QuestA-Qwen3-1.7B is a 1.7 billion parameter research checkpoint from the Qwen3 model family, developed by BytedTsinghua-SIA as part of the Direct-OPD collection. This model is specifically designed for reproducible research into post-training and reinforcement-learning-style optimization, focusing on reasoning-oriented language models. It serves as a valuable resource for studying optimization behaviors and comparing Direct-OPD checkpoints across different model families and scales.

Loading preview...

Model Overview

QuestA-Qwen3-1.7B is a 1.7 billion parameter research checkpoint from the Qwen3 model family, developed by BytedTsinghua-SIA. It is part of the Direct-OPD collection, which focuses on reproducible research for post-training and reinforcement-learning-style optimization of reasoning-oriented language models. The model is a mirrored checkpoint from ModelScope, intended to facilitate research access.

Key Capabilities & Features

  • Research Checkpoint: Specifically released for studying optimization behaviors in language models.
  • Qwen3 Family: Based on the Qwen3 architecture, providing a foundation for further research.
  • Direct-OPD Collection: Contributes to a collection designed for comparing checkpoints across various model families and parameter scales.
  • Post-training Optimization: Useful for investigating the effects of post-training and RL-style optimization methods.

Intended Use Cases

  • Studying Optimization: Ideal for analyzing post-training and RL-style optimization behaviors.
  • Comparative Research: Suitable for comparing Direct-OPD checkpoints across different model families and scales.
  • Offline Evaluation: Can be used for offline evaluation on reasoning, instruction-following, and alignment benchmarks.
  • Research Initialization: Serves as an initialization or comparison point for new research endeavors.

Limitations

  • May generate incorrect, misleading, biased, or unsafe content.
  • Full training data and hyperparameters are not fully documented in the model card.
  • Performance is sensitive to prompt format and decoding settings.
  • Research checkpoints may show regression in general instruction-following or safety while optimizing for specific objectives.