HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-001

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 23, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-001 model is an intermediate policy checkpoint from dynamic OnlineRubrics-Every GRPO training, based on the Qwen3-4B-Instruct model with 4 billion parameters and a 32768 token context length. It is distinct from static-rubric GRPO and is intended for research use only, specifically as a policy state for Phase-1 audits. This model does not have downstream medical capabilities or safety claims and is not validated for clinical decision-making.

Loading preview...

Model Overview

This model, HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-001, represents an intermediate policy checkpoint derived from dynamic OnlineRubrics-Every GRPO training. It is built upon the Qwen/Qwen3-4B-Instruct-2507 base model, featuring 4 billion parameters and a 32768 token context length, with its 'thinking' capability disabled. This particular checkpoint is a policy state specifically utilized for Phase-1 audits.

Key Characteristics

  • Training Method: Developed through dynamic OnlineRubrics-Every GRPO training, distinguishing it from static-rubric GRPO approaches.
  • Base Model: Utilizes Qwen/Qwen3-4B-Instruct-2507, a 4B parameter model.
  • Intended Use: Strictly for research purposes only; it is not designed or validated for clinical decision-making or any downstream medical applications.
  • Limitations: No medical capability or safety claims are made regarding this model.

Technical Details

The model files are provided as veRL-exported Hugging Face inference models in BF16 precision. The original_checkpoint/ directory contains the exact original FSDP parameter checkpoint along with tokenizer and configuration files, preserving the original state due to potential differences in export precision and serialization.