HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-024
The HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-024 is a 2 billion parameter Qwen3-based causal language model, fine-tuned using the GRPO reinforcement learning algorithm. This specific checkpoint is a research artifact from an experiment on reward saturation and static-rubric staleness, focusing on the medical domain. It is intended for research into policy optimization dynamics rather than direct application.
Loading preview...
Model Overview
This model, HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-024, is a 2 billion parameter policy checkpoint derived from the Qwen/Qwen3-1.7B base model. It was developed as part of an experiment by HYU-NLP-EVAL to study reward saturation and static-rubric staleness during policy optimization using the GRPO (Generalized Reinforcement Policy Optimization) algorithm.
Key Characteristics
- Base Model: Qwen3-1.7B, a causal language model.
- Training: Fine-tuned with the GRPO reinforcement learning algorithm.
- Domain Focus: Specifically trained within the RaR Medicine domain, utilizing a frozen prompt-specific initial rubric (
R0). - Research Artifact: Represents a specific checkpoint (step 24) from a larger experimental series.
- Export Format: Provided in Hugging Face Transformers format, BF16 safetensors, including model weights, configuration, tokenizer, and chat template.
Intended Use
This model is primarily a research artifact for academic study. It is designed to facilitate investigation into the dynamics of policy optimization, particularly concerning how reward functions and static rubrics behave over training steps. It is explicitly stated that medicine checkpoints are not medical devices and should not be used as a substitute for professional medical advice.