HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-003

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 26, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-003 is a 4 billion parameter Qwen3-based language model, developed by HYU-NLP-EVAL, with a 32768 token context length. This model is a research checkpoint from the static-rubric R0 experiment, specifically optimized through three static-rubric RL optimizer updates. It is designed for studying proxy-rubric staleness during policy optimization, using a frozen, prompt-specific rubric bank for training rewards. Its primary use is for research into reinforcement learning optimization techniques rather than direct application as a medical device.

Loading preview...

Overview

HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-003 is a 4 billion parameter Qwen3-based language model, specifically a research checkpoint from the static-rubric R0 experiment. This model has undergone three static-rubric RL optimizer updates, with its training rewards derived from a frozen, prompt-specific rubric bank, meaning each training prompt is scored with its own R0(x) throughout optimization. The base model for this checkpoint is Qwen/Qwen3-4B-Instruct-2507.

Key Characteristics

  • Research Focus: Primarily intended for studying proxy-rubric staleness during policy optimization, not for direct medical application.
  • Training Methodology: Utilizes a static-rubric Reinforcement Learning (RL) approach, where the reward system is based on a fixed, prompt-specific rubric bank.
  • Provenance: This specific checkpoint is pi_3 from the pilot-static-r0-100step-20260821 run, representing the third optimizer step.
  • Policy Training: Trained on a split of 256 HealthBench prompts using full-model RL parameterization.
  • Technical Specifications: Exported in BF16 dtype, with an original format of VERL FSDP v1.

Intended Use and Limitations

This model is a research checkpoint and not a medical device. It should not be used as a substitute for professional medical advice. The improvement in static-rubric reward does not inherently guarantee improvement against independent HealthBench ground truth, highlighting its experimental nature.