HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-021

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 30, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-021 is a 4 billion parameter Qwen3-based language model, fine-tuned using the GRPO method on the RaR-Medicine dataset. This model is an intermediate research checkpoint specifically developed for the medicine domain, focusing on reward-based optimization. It is designed for research into policy optimization within medical contexts, rather than clinical application.

Loading preview...

Overview

This model, HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-021, is a 4 billion parameter policy derived from Qwen/Qwen3-4B-Instruct-2507. It represents an intermediate research checkpoint from a matched static-rubric GRPO (Global Reward Policy Optimization) run, specifically after 21 global optimizer updates out of a planned 48. The model was trained using the RaR-Medicine dataset, comprising 1,500 prompts, with a focus on a rar_static_r0_only reward source within the medicine domain.

Key Characteristics

  • Architecture: Based on the Qwen3-4B-Instruct model.
  • Training Method: Utilizes the GRPO method with a static-rubric approach.
  • Domain Specificity: Fine-tuned exclusively on the RaR-Medicine dataset for medical contexts.
  • Intermediate Checkpoint: This is a research checkpoint, not a final or clinically validated model.
  • Inference: Provided as a BF16 Transformers export for efficient inference.

Intended Use Cases

  • Research in Medical NLP: Ideal for researchers exploring reward-based policy optimization and fine-tuning techniques in the medical field.
  • GRPO Experimentation: Suitable for studying the effects of GRPO updates and static-rubric reward mechanisms.
  • Comparative Analysis: Can be used to compare different stages of policy development within the specified experimental setup.

Important Note

This model is explicitly stated as an intermediate research checkpoint and not a clinical model. No medical capability or safety claims are made, and it should not be used for clinical applications.