BytedTsinghua-SIA/JustRL-Qwen3-1.7B
JustRL-Qwen3-1.7B is a 1.7 billion parameter research checkpoint from the Qwen3 model family, developed by BytedTsinghua-SIA as part of the Direct-OPD collection. It is specifically optimized using the JustRL training recipe for reproducible research on post-training and reinforcement-learning-style optimization, particularly for reasoning-oriented language models. This model is intended for studying optimization behaviors and comparing Direct-OPD checkpoints across different model families and scales.
Loading preview...
Model Overview
JustRL-Qwen3-1.7B is a 1.7 billion parameter research checkpoint within the Qwen3 model family, developed by BytedTsinghua-SIA. It is part of the Direct-OPD collection, focusing on reproducible research in post-training and reinforcement learning (RL)-style optimization for language models, specifically targeting reasoning capabilities. The model was trained using the JustRL recipe with a KL coefficient of 1.25.
Key Capabilities & Purpose
This model is primarily intended for research use, including:
- Studying Optimization Behavior: Analyzing post-training and RL-style optimization processes.
- Comparative Analysis: Comparing Direct-OPD checkpoints across various model families and parameter scales.
- Offline Evaluation: Conducting offline evaluations on reasoning, instruction-following, and alignment benchmarks.
- Research Initialization: Serving as a baseline or initialization point for further research endeavors.
Intended Use & Limitations
JustRL-Qwen3-1.7B is released under an Apache-2.0 license and is suitable for academic and research purposes. It is not intended for direct deployment in high-stakes applications (e.g., medical, legal, financial) without independent evaluation and robust safeguards. Users should be aware that the model may generate incorrect, biased, or unsafe content, and its performance can be sensitive to prompt formats and decoding settings. Full training data and hyperparameter details are not fully documented in this model card, and research checkpoints may sometimes regress on general instruction-following while improving on targeted objectives.