HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-006
HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-006 is a 1.7 billion parameter Qwen3-based causal language model developed by HYU-NLP-EVAL. This model is a research artifact from an experiment on reward saturation and static-rubric staleness during policy optimization, specifically fine-tuned for the RaR Medicine domain. It is intended for research into policy optimization and not for medical advice or general-purpose applications.
Loading preview...
Overview
This model, HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-006, is a 1.7 billion parameter policy checkpoint derived from the Qwen/Qwen3-1.7B base model. It was developed by HYU-NLP-EVAL as part of an experiment investigating static-rubric discriminability and reward saturation in policy optimization.
Key Characteristics
- Base Model: Qwen3-1.7B, with a base revision of
70d244cc86ccca08cf5af4e1e306ecf908b1ad5e. - Training Method: Utilizes the GRPO (Generalized Policy Optimization) reinforcement learning algorithm.
- Reward Mechanism: Trained with a frozen, prompt-specific initial rubric (R0).
- Domain Specificity: Specifically fine-tuned for the RaR Medicine domain, covering ten planned audit points through step 48 of the training process.
- Export Format: Provided in Hugging Face Transformers format, using BF16 safetensors, including model weights, configuration, tokenizer, and chat template.
Intended Use
This checkpoint is strictly a research artifact designed for studying reward saturation and the staleness of static rubrics during policy optimization. It is explicitly stated that these medicine checkpoints are not medical devices and should not be used as a substitute for professional medical advice. Developers should use this model for academic research purposes related to RL policy optimization in domain-specific contexts.