BytedTsinghua-SIA/JustRL-R1-7B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

JustRL-R1-7B is a 7 billion parameter research checkpoint from the BytedTsinghua-SIA Direct-OPD collection, based on the R1 model family. It is specifically designed for reproducible research into post-training and reinforcement learning-style optimization for reasoning-oriented language models. This model is intended for studying optimization behaviors and comparing Direct-OPD checkpoints, serving as an initialization point for further research.

Loading preview...

JustRL-R1-7B: A Research Checkpoint for RL-Style Optimization

JustRL-R1-7B is a 7 billion parameter research checkpoint developed by BytedTsinghua-SIA, part of the Direct-OPD collection. This model is specifically released to facilitate reproducible research in the domain of post-training and reinforcement learning (RL)-style optimization for language models, with a particular focus on reasoning capabilities.

Key Characteristics

  • Model Family: R1, with 7 billion parameters.
  • Training Method: Utilizes the JustRL training recipe with an adaptive KL setting.
  • Provenance: Mirrored from ModelScope to Hugging Face to enhance accessibility within the Direct-OPD collection.
  • Research Focus: Designed for studying the behavior of post-training and RL-style optimization.

Intended Use Cases

This model is primarily intended for research purposes, including:

  • Optimization Research: Investigating post-training and RL-style optimization behaviors.
  • Comparative Analysis: Comparing Direct-OPD checkpoints across different model families and parameter scales.
  • Offline Evaluation: Conducting offline evaluations on benchmarks related to reasoning, instruction-following, and alignment.
  • Research Initialization: Serving as a baseline or initialization point for further research and development.

Important Note: This model is not intended for direct deployment in safety-critical or high-stakes applications without thorough independent evaluation and robust safeguards. Users are advised to conduct their own evaluations on relevant tasks and safety criteria.