HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-013

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-013 model is an intermediate policy checkpoint derived from dynamic OnlineRubrics-Every GRPO training, distinct from static-rubric GRPO. Based on the Qwen3-4B-Instruct model with 4 billion parameters and a 32768 token context length, this version has thinking disabled. It represents a historical policy state used in a Phase-1 audit, intended for research use only without claims of medical capability or safety. This model is specifically designed for research in the context of OnlineRubrics RaR-Medicine.

Loading preview...

Model Overview

This model, HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-013, is an intermediate policy checkpoint resulting from dynamic OnlineRubrics-Every GRPO training. It is built upon the Qwen/Qwen3-4B-Instruct-2507 base model, featuring 4 billion parameters and a 32768 token context length, with its 'thinking' capability explicitly disabled.

Key Characteristics

  • Training Origin: Derived from dynamic OnlineRubrics-Every GRPO training, which differs from static-rubric GRPO methods.
  • Base Model: Utilizes Qwen/Qwen3-4B-Instruct-2507 as its foundation.
  • Policy State: Represents a specific historical policy state that was employed during a Phase-1 audit.
  • Research Focus: Explicitly designated for research use only, with no claims regarding downstream medical capabilities or safety. It has not been validated for clinical decision-making.

Technical Details

  • The model files are veRL-exported Hugging Face inference models in BF16 format.
  • The original_checkpoint/ directory contains the exact original FSDP parameter checkpoint along with tokenizer and configuration files.
  • Crucially, optimizer state, training data, responses, rubrics, infrastructure configuration, and credentials are not included in this release.

Intended Use

  • Research: Primarily intended for research purposes within the OnlineRubrics RaR-Medicine context.
  • Historical Analysis: Useful for understanding the policy state at a specific point in the training process for audit or analysis.