HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-013
HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-013 is a 1.7 billion parameter Qwen3-based causal language model, fine-tuned using the GRPO reinforcement learning algorithm. This specific checkpoint is a research artifact from an experiment on reward saturation and static-rubric staleness, focusing on the medical domain. It is intended for research into policy optimization dynamics rather than direct application.
Loading preview...
Overview
This repository hosts a policy checkpoint, qwen3-1.7b-rar-medicine-static-r0-step-013, derived from the Qwen/Qwen3-1.7B base model. It is part of an experiment investigating reward saturation and static-rubric staleness during policy optimization using the GRPO (Generalized Policy Optimization) algorithm. The model has 1.7 billion parameters and a context length of 32768 tokens.
Key Characteristics
- Base Model: Qwen3-1.7B
- RL Algorithm: GRPO (Generalized Policy Optimization)
- Reward Function: Frozen prompt-specific initial rubric (
R0) - Domain: RaR Medicine, indicating its training focus on medical-related data.
- Checkpoint Contents: Includes model weights, configuration, tokenizer, and chat template, exported in BF16 safetensors format.
- Research Focus: Specifically designed as a research artifact to study the dynamics of policy optimization, particularly how reward functions and rubrics behave over training steps.
Intended Use
This model is explicitly designated as a research artifact. Its primary purpose is to facilitate studies on reward saturation and the staleness of static rubrics during policy optimization. It is crucial to note that these medicine checkpoints are not medical devices and should not be used as a substitute for professional medical advice. They are for experimental analysis within a research context.