HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-032
HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-032 is a 1.7 billion parameter causal language model based on Qwen3, developed by HYU-NLP-EVAL. This specific checkpoint is a research artifact from a static-rubric discriminability-horizon experiment, fine-tuned using the GRPO algorithm with a frozen prompt-specific initial rubric (R0) in the medical domain. It is intended for studying reward saturation and static-rubric staleness during policy optimization, rather than direct application.
Loading preview...
Model Overview
This repository hosts a specific policy checkpoint, qwen3-1.7b-rar-medicine-static-r0-step-032, derived from the Qwen3-1.7B base model. Developed by HYU-NLP-EVAL, this model is a research artifact from an experiment focused on static-rubric discriminability-horizon within the medical domain. It utilizes the GRPO (Generalized Reinforcement Policy Optimization) algorithm and was trained with a frozen prompt-specific initial rubric (R0).
Key Characteristics
- Base Model: Qwen/Qwen3-1.7B, revision
70d244cc86ccca08cf5af4e1e306ecf908b1ad5e. - Training Algorithm: GRPO (Generalized Reinforcement Policy Optimization).
- Reward Mechanism: Uses a frozen prompt-specific initial rubric (R0).
- Domain: Specifically trained within the RaR Medicine domain.
- Format: Exported in Hugging Face Transformers format, using BF16 safetensors.
- Contents: Includes model weights, configuration, tokenizer, and chat template.
- Exclusions: Optimizer, scheduler, trainer state, rollouts, rubrics, and evaluation data are not included.
Intended Use
This checkpoint is primarily intended as a research artifact for academic study. Its main purpose is to facilitate research into reward saturation and the staleness of static rubrics during policy optimization processes. It is explicitly stated that these medicine checkpoints are not medical devices and should not be used as a substitute for professional medical advice.