HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-016
HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-016 is a 2 billion parameter Qwen3-1.7B policy checkpoint, developed by HYU-NLP-EVAL, derived from the Qwen/Qwen3-1.7B base model. This model is a research artifact from an experiment studying reward saturation and static-rubric staleness during policy optimization, specifically within the RaR Medicine domain. It is fine-tuned using the GRPO RL algorithm with a frozen prompt-specific initial rubric (R0). This checkpoint is intended for research into policy optimization dynamics rather than direct application.
Loading preview...
Model Overview
HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-016 is a 2 billion parameter policy checkpoint based on the Qwen/Qwen3-1.7B model. It is a research artifact from an experiment conducted by HYU-NLP-EVAL, focusing on the dynamics of reward saturation and static-rubric staleness during policy optimization.
Key Characteristics
- Base Model: Qwen/Qwen3-1.7B
- RL Algorithm: Trained using GRPO (Generalized Reward Policy Optimization).
- Reward Mechanism: Utilizes a frozen prompt-specific initial rubric (
R0) for training. - Domain: Specifically trained within the RaR Medicine domain, with this checkpoint representing step 16 of the optimization process.
- Export Format: Provided in Hugging Face Transformers format, BF16 safetensors, including model weights, configuration, tokenizer, and chat template.
Intended Use
This model is primarily intended for academic and research purposes, specifically for:
- Studying reward saturation in reinforcement learning policies.
- Investigating static-rubric staleness during policy optimization.
Important Note: Medicine checkpoints like this one are research artifacts and are not medical devices. They must not be used as a substitute for professional medical advice or in any clinical application.