HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-002
The HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-002 is a 4 billion parameter language model, based on the Qwen3-4B-Instruct architecture, developed as an intermediate policy from dynamic OnlineRubrics-Every GRPO training. This model is specifically designed for research use in the medical domain, serving as a policy state for Phase-1 audits. It is distinct from static-rubric GRPO models and is not validated for clinical decision-making, with thinking capabilities disabled.
Loading preview...
HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-002 Overview
This model is an intermediate policy checkpoint, derived from dynamic OnlineRubrics-Every GRPO training, and is distinct from static-rubric GRPO methods. Built upon the Qwen/Qwen3-4B-Instruct-2507 base model, it features 4 billion parameters and has its 'thinking' capabilities explicitly disabled. This specific checkpoint represents a policy state utilized during a Phase-1 audit.
Key Characteristics
- Architecture: Based on the Qwen3-4B-Instruct model.
- Training Method: Developed via dynamic OnlineRubrics-Every GRPO, differing from static-rubric GRPO.
- Parameter Count: 4 billion parameters.
- Context Length: Supports a context length of 32768 tokens.
- Purpose: Primarily intended for research use as an intermediate policy state.
- Limitations: No downstream medical capability or safety claims are made; it is not validated for clinical decision-making.
Intended Use Cases
- Research: Ideal for academic and research purposes, particularly in understanding dynamic GRPO training and policy states.
- Auditing: Suitable for use in Phase-1 audit processes where this specific policy state is relevant.
Root files are provided as a veRL-exported Hugging Face inference model in BF16 precision. The original FSDP parameter checkpoint and tokenizer/configuration files are preserved for exact replication, as export precision and serialization may differ.