HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-003
HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-003 is a 4 billion parameter Qwen3-based language model, developed by HYU-NLP-EVAL, with a 32768 token context length. This model is a research checkpoint from the static-rubric R0 experiment, specifically optimized through three static-rubric RL optimizer updates. It is designed for studying proxy-rubric staleness during policy optimization, using a frozen, prompt-specific rubric bank for training rewards. Its primary use is for research into reinforcement learning optimization techniques rather than direct application as a medical device.
Loading preview...
Overview
HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-003 is a 4 billion parameter Qwen3-based language model, specifically a research checkpoint from the static-rubric R0 experiment. This model has undergone three static-rubric RL optimizer updates, with its training rewards derived from a frozen, prompt-specific rubric bank, meaning each training prompt is scored with its own R0(x) throughout optimization. The base model for this checkpoint is Qwen/Qwen3-4B-Instruct-2507.
Key Characteristics
- Research Focus: Primarily intended for studying proxy-rubric staleness during policy optimization, not for direct medical application.
- Training Methodology: Utilizes a static-rubric Reinforcement Learning (RL) approach, where the reward system is based on a fixed, prompt-specific rubric bank.
- Provenance: This specific checkpoint is
pi_3from thepilot-static-r0-100step-20260821run, representing the third optimizer step. - Policy Training: Trained on a split of 256 HealthBench prompts using full-model RL parameterization.
- Technical Specifications: Exported in BF16 dtype, with an original format of VERL FSDP v1.
Intended Use and Limitations
This model is a research checkpoint and not a medical device. It should not be used as a substitute for professional medical advice. The improvement in static-rubric reward does not inherently guarantee improvement against independent HealthBench ground truth, highlighting its experimental nature.