HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-046

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 30, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-046 model is an intermediate policy derived from dynamic OnlineRubrics-Every GRPO training, based on the Qwen3-4B-Instruct-2507 architecture with 4 billion parameters and a 32768 token context length. This model is specifically a research checkpoint from a Phase-1 audit, distinct from static-rubric GRPO, and is intended for research use only without medical capability or safety claims. It represents a fine-tuned policy state for specific research into online rubric-based reinforcement learning in a medical context.

Loading preview...

Model Overview

HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-046 is an intermediate policy model developed by HYU-NLP-EVAL. It is a 4-billion parameter model built upon the Qwen/Qwen3-4B-Instruct-2507 base, featuring a context length of 32768 tokens. This specific checkpoint, step 46, seed 11, originates from dynamic OnlineRubrics-Every GRPO training, a method distinct from static-rubric GRPO.

Key Characteristics

  • Training Origin: Derived from dynamic OnlineRubrics-Every GRPO training, focusing on policy development.
  • Base Model: Utilizes Qwen/Qwen3-4B-Instruct-2507 as its foundation, with 'thinking disabled' during its specific training phase.
  • Research Focus: This model is explicitly a policy state from a Phase-1 audit, intended solely for research purposes.
  • No Medical Claims: The developers explicitly state that no downstream medical capability or safety claims are made, and it is not validated for clinical decision-making.

Intended Use

This model is primarily for:

  • Research Use: Ideal for researchers studying dynamic online rubric-based reinforcement learning and policy development in specialized domains.
  • Auditing and Analysis: Useful for analyzing intermediate policy states within the OnlineRubrics-Every GRPO training framework.

It is crucial to note that this model is not for clinical decision-making or any application requiring validated medical capabilities.