HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-048
The HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-048 is a 1.7 billion parameter Qwen3-based language model, fine-tuned using the GRPO reinforcement learning algorithm. This specific checkpoint, from step 48, is optimized for research into reward saturation and static-rubric staleness within the medical domain. It serves as a research artifact for studying policy optimization dynamics rather than a general-purpose medical advice tool.
Loading preview...
Overview
This model, HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-048, is a 1.7 billion parameter policy checkpoint derived from the Qwen/Qwen3-1.7B base model. It was developed by HYU-NLP-EVAL as part of an experiment investigating reward saturation and static-rubric staleness in policy optimization using the GRPO (Generalized Reinforcement Policy Optimization) algorithm.
Key Characteristics
- Base Model: Qwen3-1.7B, a causal language model.
- Training Method: Fine-tuned using the GRPO reinforcement learning algorithm.
- Reward Mechanism: Utilizes a frozen, prompt-specific initial rubric (
R0). - Domain Focus: Specifically trained within the Medicine domain, covering ten planned audit points through step 48.
- Format: Exported in Hugging Face Transformers format, using BF16 safetensors.
- Contents: Includes model weights, configuration, tokenizer, and chat template.
Intended Use
This checkpoint is primarily a research artifact.
- Study Focus: Designed for studying reward saturation and the staleness of static rubrics during policy optimization.
- Limitations: It is explicitly stated that these medicine checkpoints are not medical devices and should not be used as a substitute for professional medical advice. Its purpose is purely for experimental analysis within the NLP research community.