HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-001
The HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-001 model is an intermediate policy checkpoint from dynamic OnlineRubrics-Every GRPO training, based on the Qwen3-4B-Instruct model with 4 billion parameters and a 32768 token context length. It is distinct from static-rubric GRPO and is intended for research use only, specifically as a policy state for Phase-1 audits. This model does not have downstream medical capabilities or safety claims and is not validated for clinical decision-making.
Loading preview...
Model Overview
This model, HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-001, represents an intermediate policy checkpoint derived from dynamic OnlineRubrics-Every GRPO training. It is built upon the Qwen/Qwen3-4B-Instruct-2507 base model, featuring 4 billion parameters and a 32768 token context length, with its 'thinking' capability disabled. This particular checkpoint is a policy state specifically utilized for Phase-1 audits.
Key Characteristics
- Training Method: Developed through dynamic OnlineRubrics-Every GRPO training, distinguishing it from static-rubric GRPO approaches.
- Base Model: Utilizes Qwen/Qwen3-4B-Instruct-2507, a 4B parameter model.
- Intended Use: Strictly for research purposes only; it is not designed or validated for clinical decision-making or any downstream medical applications.
- Limitations: No medical capability or safety claims are made regarding this model.
Technical Details
The model files are provided as veRL-exported Hugging Face inference models in BF16 precision. The original_checkpoint/ directory contains the exact original FSDP parameter checkpoint along with tokenizer and configuration files, preserving the original state due to potential differences in export precision and serialization.