BytedTsinghua-SIA/JustRL-Qwen3-4B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

JustRL-Qwen3-4B is a 4 billion parameter research checkpoint from the Qwen3 family, developed by BytedTsinghua-SIA as part of the Direct-OPD collection. It is specifically optimized for reproducible research on post-training and reinforcement-learning-style optimization for reasoning-oriented language models. This model is intended for studying optimization behaviors and serving as an initialization point for further research rather than direct deployment.

Loading preview...

Model Overview

JustRL-Qwen3-4B is a 4 billion parameter research checkpoint from the Qwen3 model family, developed by BytedTsinghua-SIA. It is part of the Direct-OPD collection, focusing on reproducible research in post-training and reinforcement-learning-style optimization for reasoning-oriented language models. The model was mirrored from ModelScope to Hugging Face to facilitate access for the research community.

Key Characteristics

  • Model Family: Qwen3
  • Parameter Scale: 4 Billion
  • Training Recipe: JustRL, with a KL coefficient of 2.
  • Provenance: A Direct-OPD checkpoint, intended for studying optimization behaviors.

Intended Use Cases

This model is primarily designed for research purposes, including:

  • Studying post-training and RL-style optimization behaviors.
  • Comparing Direct-OPD checkpoints across different model families and parameter scales.
  • Running offline evaluations on reasoning, instruction-following, and alignment benchmarks.
  • Serving as an initialization or comparison point for further research.

Limitations

As a research checkpoint, JustRL-Qwen3-4B may generate incorrect, misleading, or biased content. The full training dataset and hyperparameter schedule are not fully documented in the model card. It is not intended for direct deployment in safety-critical or high-stakes settings without independent evaluation and safeguards.