HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-009
HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-009 is a 4 billion parameter Qwen3-based causal language model, fine-tuned by HYU-NLP-EVAL using the GRPO method on the RaR-Medicine dataset. This model is an intermediate research checkpoint specifically trained for medicine-related tasks, focusing on static-rubric reward optimization. It is designed for research in medical language processing, with a context length of 32768 tokens.
Loading preview...
HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-009
This model is an intermediate research checkpoint from the HYU-NLP-EVAL group, based on the Qwen3-4B-Instruct architecture. It represents the policy after 9 global optimizer updates of a matched static-rubric GRPO (Global Reward Policy Optimization) run, specifically trained on the RaR-Medicine dataset.
Key Characteristics
- Architecture: Qwen3-4B-Instruct, a 4 billion parameter causal language model.
- Training Method: Fine-tuned using the GRPO method with a
static_r0_matchedapproach. - Reward Source: Utilizes
rar_static_r0_onlyfor reward signals. - Domain Specialization: Trained on 1,500 prompts from the RaR-Medicine dataset, focusing on medical language processing.
- Context Length: Supports a context length of 32768 tokens.
- Research Focus: This is a research checkpoint, not intended for clinical use, and no medical capability or safety claims are made.
Intended Use
This model is suitable for researchers and developers exploring:
- The application of GRPO methods in domain-specific language models.
- Performance of Qwen3-4B-Instruct in medical contexts.
- Further fine-tuning or analysis within the medical NLP research domain.
It is important to note that this is an ongoing research project, and the model is provided as an intermediate checkpoint.