HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-000
HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-000 is a 2 billion parameter Qwen3-1.7B based causal language model, fine-tuned using the GRPO algorithm with a frozen prompt-specific initial rubric (R0) reward. This model is a research artifact from an experiment studying reward saturation and static-rubric staleness in policy optimization, specifically within the RaR Medicine domain. It is intended for research into RL policy optimization rather than direct application.
Loading preview...
Overview
This model, HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-000, is a research artifact derived from the Qwen3-1.7B base policy. It represents a policy checkpoint from a static-rubric discriminability-horizon experiment, specifically focusing on the RaR Medicine domain. The model was fine-tuned using the GRPO (Generalized Policy Optimization) algorithm, with training guided by a frozen prompt-specific initial rubric (R0) as the reward signal.
Key Characteristics
- Base Model: Qwen/Qwen3-1.7B, a 2 billion parameter causal language model.
- Fine-tuning Method: GRPO (Generalized Policy Optimization) algorithm.
- Reward Mechanism: Utilizes a frozen prompt-specific initial rubric (R0) for training.
- Domain Focus: Part of an experiment in the RaR Medicine domain, covering ten planned audit points through step 48.
- Export Format: Hugging Face Transformers, BF16 safetensors, including model weights, configuration, tokenizer, and chat template.
Intended Use
This checkpoint is primarily intended as a research artifact for studying reward saturation and the staleness of static rubrics during policy optimization. It is explicitly stated that medicine checkpoints are not medical devices and should not be used as a substitute for professional medical advice. Its value lies in contributing to the understanding of RL policy development rather than direct deployment in medical or other applications.