HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-009

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 30, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-009 is a 4 billion parameter Qwen3-based causal language model, fine-tuned by HYU-NLP-EVAL using the GRPO method on the RaR-Medicine dataset. This model is an intermediate research checkpoint specifically trained for medicine-related tasks, focusing on static-rubric reward optimization. It is designed for research in medical language processing, with a context length of 32768 tokens.

Loading preview...

HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-009

This model is an intermediate research checkpoint from the HYU-NLP-EVAL group, based on the Qwen3-4B-Instruct architecture. It represents the policy after 9 global optimizer updates of a matched static-rubric GRPO (Global Reward Policy Optimization) run, specifically trained on the RaR-Medicine dataset.

Key Characteristics

  • Architecture: Qwen3-4B-Instruct, a 4 billion parameter causal language model.
  • Training Method: Fine-tuned using the GRPO method with a static_r0_matched approach.
  • Reward Source: Utilizes rar_static_r0_only for reward signals.
  • Domain Specialization: Trained on 1,500 prompts from the RaR-Medicine dataset, focusing on medical language processing.
  • Context Length: Supports a context length of 32768 tokens.
  • Research Focus: This is a research checkpoint, not intended for clinical use, and no medical capability or safety claims are made.

Intended Use

This model is suitable for researchers and developers exploring:

  • The application of GRPO methods in domain-specific language models.
  • Performance of Qwen3-4B-Instruct in medical contexts.
  • Further fine-tuning or analysis within the medical NLP research domain.

It is important to note that this is an ongoing research project, and the model is provided as an intermediate checkpoint.