HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-015
HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-015 is a 4 billion parameter policy model based on Qwen3-4B-Instruct, fine-tuned using the GRPO method with a static-rubric reward source on the RaR-Medicine dataset. This intermediate research checkpoint, developed by HYU-NLP-EVAL, is specifically trained on 1,500 medical prompts. It represents the 15th global optimizer update in a matched static-rubric GRPO run, focusing on medical domain applications.
Loading preview...
Overview
This model, HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-015, is an intermediate research checkpoint from a GRPO (Global Reward Policy Optimization) experiment. It is a 4 billion parameter policy derived from Qwen/Qwen3-4B-Instruct-2507, fine-tuned using a static-rubric reward source (rar_static_r0_only) within the medical domain.
Key Characteristics
- Base Model: Qwen3-4B-Instruct-2507.
- Training Method: GRPO with
static_r0_matchedapproach. - Domain: Medicine, trained on 1,500 prompts from the RaR-Medicine dataset.
- Development Stage: Represents the 15th global optimizer update in a planned 48-step run.
- Configuration: Uses a global prompt batch of 96, 16 rollouts per prompt, and a learning rate of 5e-06.
- Inference: Provided as a BF16 Transformers export for efficient inference.
Intended Use and Limitations
This model is a research checkpoint and is explicitly not a clinical model. No medical capability or safety claims are made. It is suitable for researchers exploring GRPO methods and their application in specialized domains like medicine, particularly for understanding the effects of static-rubric reward optimization.