HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-013
The HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-013 model is an intermediate policy checkpoint derived from dynamic OnlineRubrics-Every GRPO training, distinct from static-rubric GRPO. Based on the Qwen3-4B-Instruct model with 4 billion parameters and a 32768 token context length, this version has thinking disabled. It represents a historical policy state used in a Phase-1 audit, intended for research use only without claims of medical capability or safety. This model is specifically designed for research in the context of OnlineRubrics RaR-Medicine.
Loading preview...
Model Overview
This model, HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-013, is an intermediate policy checkpoint resulting from dynamic OnlineRubrics-Every GRPO training. It is built upon the Qwen/Qwen3-4B-Instruct-2507 base model, featuring 4 billion parameters and a 32768 token context length, with its 'thinking' capability explicitly disabled.
Key Characteristics
- Training Origin: Derived from dynamic OnlineRubrics-Every GRPO training, which differs from static-rubric GRPO methods.
- Base Model: Utilizes
Qwen/Qwen3-4B-Instruct-2507as its foundation. - Policy State: Represents a specific historical policy state that was employed during a Phase-1 audit.
- Research Focus: Explicitly designated for research use only, with no claims regarding downstream medical capabilities or safety. It has not been validated for clinical decision-making.
Technical Details
- The model files are veRL-exported Hugging Face inference models in BF16 format.
- The
original_checkpoint/directory contains the exact original FSDP parameter checkpoint along with tokenizer and configuration files. - Crucially, optimizer state, training data, responses, rubrics, infrastructure configuration, and credentials are not included in this release.
Intended Use
- Research: Primarily intended for research purposes within the OnlineRubrics RaR-Medicine context.
- Historical Analysis: Useful for understanding the policy state at a specific point in the training process for audit or analysis.