HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-000

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 26, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-000 is a 4 billion parameter Qwen3-based causal language model. This model represents an initialization checkpoint (step 0) from a research experiment focused on proxy-rubric staleness during policy optimization. It is specifically derived from a Qwen/Qwen3-4B-Instruct base model and is part of a static-rubric R0 experiment, where training rewards use a frozen, prompt-specific rubric bank for HealthBench prompts. Its primary purpose is for research into RL optimization dynamics rather than direct application.

Loading preview...

Model Overview

This model, HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-000, is an initialization checkpoint (step 0) of a 4 billion parameter Qwen3-based causal language model. It is part of a research experiment designed to study proxy-rubric staleness during policy optimization, specifically within the context of HealthBench prompts.

Key Characteristics

  • Base Model: Derived from Qwen/Qwen3-4B-Instruct-2507.
  • Research Focus: This checkpoint is from a "static-rubric R0" experiment, meaning training rewards utilize a frozen, prompt-specific rubric bank (R0(x)) throughout the optimization process.
  • Training Context: The policy training split involved 256 HealthBench prompts, with full-model Reinforcement Learning (RL) parameterization.
  • Export Format: Exported in BF16 data type, originating from VERL FSDP v1 with FP32 state dict.

Intended Use and Limitations

  • Research Tool: This model is strictly a research checkpoint for investigating proxy-rubric staleness in RL policy optimization.
  • Not for Medical Use: It is explicitly stated not to be a medical device and should not be used as a substitute for professional medical advice.
  • Evaluation Caveat: Improvements in static-rubric reward do not inherently guarantee improvement against independent HealthBench ground truth.