HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-024

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 30, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-024 is a 4 billion parameter policy model based on Qwen3-4B-Instruct, fine-tuned using the GRPO method on the RaR-Medicine dataset. This model is an intermediate research checkpoint focused on medical domain applications, specifically trained with a static-rubric reward function. It is designed for research into reinforcement learning policies in medicine, rather than clinical deployment.

Loading preview...

Model Overview

This model, HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-024, is a 4 billion parameter policy derived from Qwen/Qwen3-4B-Instruct-2507. It represents an intermediate research checkpoint from a matched static-rubric GRPO (Global Reinforcement Policy Optimization) run, specifically at step 24 of 48 planned updates.

Key Characteristics

  • Base Model: Qwen3-4B-Instruct-2507, a 4 billion parameter causal language model.
  • Fine-tuning Method: Utilizes the GRPO method with a static_r0_matched approach.
  • Reward Source: rar_static_r0_only, indicating a reward function based on a static rubric.
  • Domain: Specialized for the Medicine domain, trained on 1,500 prompts from the RaR-Medicine dataset.
  • Training Configuration: Features a learning rate of 5e-06, with a GRPO global prompt batch of 96 and 16 rollouts per prompt. Thinking capabilities were disabled during training.
  • Intermediate Checkpoint: This is a research artifact and not intended for clinical use, with no medical capability or safety claims made.

Intended Use

This model is suitable for researchers exploring reinforcement learning policies in the medical domain, particularly those interested in the effects of static-rubric reward functions and GRPO training. It provides a specific snapshot of a policy's evolution during a controlled research experiment.