HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-027
HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-027 is a 4 billion parameter language model, an intermediate policy checkpoint from dynamic OnlineRubrics-Every GRPO training, distinct from static-rubric GRPO. Based on Qwen/Qwen3-4B-Instruct-2507 with thinking disabled, this model represents a historical policy state used by the Phase-1 audit. It is intended for research use only, with no medical capability or safety claims, and is not validated for clinical decision-making.
Loading preview...
Model Overview
This model, HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-027, is a 4 billion parameter language model derived from Qwen/Qwen3-4B-Instruct-2507. It represents an intermediate policy checkpoint from a dynamic OnlineRubrics-Every GRPO training process, which differs from static-rubric GRPO methods. Notably, the 'thinking' capability is disabled in this specific iteration.
Key Characteristics
- Training Origin: An intermediate policy from dynamic OnlineRubrics-Every GRPO training.
- Base Model: Built upon Qwen/Qwen3-4B-Instruct-2507.
- Historical Checkpoint: This specific checkpoint served as a historical policy state during a Phase-1 audit.
- Research Use Only: The model is explicitly for research purposes; no medical capability or safety claims are made, and it is not validated for clinical decision-making.
Technical Details
- Parameter Count: 4 billion parameters.
- Context Length: Supports a context length of 32768 tokens.
- Precision: Root files are provided as veRL-exported Hugging Face inference model in BF16 format.
Important Note
The original_checkpoint/ directory contains the exact original FSDP parameter checkpoint along with tokenizer and configuration files. Optimizer state, training data, responses, rubrics, infrastructure configuration, and credentials are not included in this release.