HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-042

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 19, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-042 is a 4 billion parameter policy model based on Qwen3-4B-Instruct, fine-tuned using the GRPO method on the RaR-Medicine dataset. This model is an intermediate research checkpoint, specifically optimized for medical domain tasks through static-rubric reinforcement learning. It is designed for research into policy optimization within the medical field, rather than clinical application. The model has a context length of 32768 tokens and was trained with a specific focus on medical prompts.

Loading preview...

Overview

This model, HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-042, is a 4 billion parameter policy derived from the Qwen/Qwen3-4B-Instruct-2507 base model. It represents an intermediate research checkpoint, specifically step 42 out of a planned 48 global optimizer updates, from a matched static-rubric GRPO (Generative Reinforcement Policy Optimization) run.

Key Characteristics

  • Methodology: Utilizes static_r0_matched GRPO, focusing on a static rubric for reward signals.
  • Domain Specialization: Fine-tuned on the RaR-Medicine dataset, comprising 1,500 medical prompts.
  • Base Model: Built upon Qwen/Qwen3-4B-Instruct-2507 with a base revision cdbee75f17c01a7cc42f958dc650907174af0554.
  • Training Details: Trained with a learning rate of 5e-06, using a global prompt batch of 96 and 16 rollouts per prompt.
  • Context Length: Supports a context length of 32768 tokens.

Intended Use and Limitations

This model is explicitly an intermediate research checkpoint and is not intended for clinical use. No medical capability or safety claims are made. It is designed for research purposes related to policy optimization in the medical domain, particularly for exploring the effects of static-rubric GRPO on medical text generation. The original checkpoint files are available in the original_checkpoint/ directory for detailed analysis.