HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-000
HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-000 is a 4 billion parameter Qwen3-based causal language model. This model represents an initialization checkpoint (step 0) from a research experiment focused on proxy-rubric staleness during policy optimization. It is specifically derived from a Qwen/Qwen3-4B-Instruct base model and is part of a static-rubric R0 experiment, where training rewards use a frozen, prompt-specific rubric bank for HealthBench prompts. Its primary purpose is for research into RL optimization dynamics rather than direct application.
Loading preview...
Model Overview
This model, HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-000, is an initialization checkpoint (step 0) of a 4 billion parameter Qwen3-based causal language model. It is part of a research experiment designed to study proxy-rubric staleness during policy optimization, specifically within the context of HealthBench prompts.
Key Characteristics
- Base Model: Derived from
Qwen/Qwen3-4B-Instruct-2507. - Research Focus: This checkpoint is from a "static-rubric R0" experiment, meaning training rewards utilize a frozen, prompt-specific rubric bank (
R0(x)) throughout the optimization process. - Training Context: The policy training split involved 256 HealthBench prompts, with full-model Reinforcement Learning (RL) parameterization.
- Export Format: Exported in BF16 data type, originating from VERL FSDP v1 with FP32 state dict.
Intended Use and Limitations
- Research Tool: This model is strictly a research checkpoint for investigating proxy-rubric staleness in RL policy optimization.
- Not for Medical Use: It is explicitly stated not to be a medical device and should not be used as a substitute for professional medical advice.
- Evaluation Caveat: Improvements in static-rubric reward do not inherently guarantee improvement against independent HealthBench ground truth.