HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-039

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 9, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-039 model is a 4 billion parameter policy checkpoint derived from the Qwen3-4B-Instruct-2507 base model, specifically trained using dynamic OnlineRubrics-Every GRPO. This model is an intermediate policy state from a dynamic training process, distinct from static-rubric GRPO, and was used in a Phase-1 audit. It is intended for research use only, without any downstream medical capability or safety claims, and is not validated for clinical decision-making.

Loading preview...

Model Overview

HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-039 is a 4 billion parameter language model, representing an intermediate policy checkpoint from a dynamic training process. It is based on the Qwen/Qwen3-4B-Instruct-2507 model, with thinking capabilities disabled during its development. This specific checkpoint was generated during step 39 with seed 11 of an OnlineRubrics-Every GRPO training regimen, which differs from static-rubric GRPO methods.

Key Characteristics

  • Base Model: Qwen/Qwen3-4B-Instruct-2507.
  • Training Method: Dynamic OnlineRubrics-Every GRPO, indicating a training approach where rubrics are dynamically generated.
  • Purpose: This checkpoint serves as a historical policy state, specifically used for a Phase-1 audit.
  • Format: The root files are veRL-exported Hugging Face inference models in BF16 precision.

Important Considerations

  • Research Use Only: The model is strictly for research purposes. No claims are made regarding its medical capabilities or safety.
  • No Clinical Validation: It has not been validated for use in clinical decision-making.
  • Checkpoint Details: The original_checkpoint/ directory contains the exact original FSDP parameter checkpoint and tokenizer/configuration files, preserving the original state due to potential precision/serialization differences during export. Optimizer state, training data, responses, rubrics, infrastructure configuration, and credentials are not included in this release.