HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-050
HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-050 is a 4 billion parameter Qwen3-based causal language model. This research checkpoint is specifically optimized for health-related tasks, having undergone 50 static-rubric RL optimizer updates on HealthBench prompts. It is designed for studying proxy-rubric staleness during policy optimization, offering insights into reward model behavior in specialized domains. The model is intended for research purposes in natural language processing and health AI, not for medical advice.
Loading preview...
Overview
This model, HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-050, is a 4-billion parameter checkpoint derived from the Qwen/Qwen3-4B-Instruct-2507 base model. It represents pi_50, the state after 50 updates using a static-rubric Reinforcement Learning (RL) optimizer. The training utilized a frozen, prompt-specific rubric bank, meaning each training prompt was scored with its own R0(x) throughout the optimization process.
Key Characteristics
- Base Model: Qwen3-4B-Instruct-2507.
- Optimization: Full-model RL with 50 static-rubric optimizer steps.
- Training Data: Policy training was conducted on 256 HealthBench prompts.
- Purpose: Primarily a research checkpoint for investigating proxy-rubric staleness during policy optimization, particularly in health-related contexts.
Intended Use and Limitations
This model is a research checkpoint and is explicitly not a medical device. It must not be used as a substitute for professional medical advice. The improvement observed through static-rubric reward does not inherently guarantee improvement against independent HealthBench ground truth, highlighting its research-oriented nature for studying reward model dynamics rather than direct application.