HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-009

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 3, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-009 is a 1.7 billion parameter causal language model based on Qwen3, developed by HYU-NLP-EVAL. This specific checkpoint is a research artifact from a static-rubric discriminability-horizon experiment, fine-tuned using the GRPO algorithm with a frozen prompt-specific initial rubric (R0) in the medical domain. It is intended for studying reward saturation and static-rubric staleness during policy optimization, rather than direct application.

Loading preview...

Model Overview

This model, HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-009, is a specialized policy checkpoint derived from the Qwen/Qwen3-1.7B base model. It is part of a research experiment focused on understanding the dynamics of reward saturation and the staleness of static rubrics during policy optimization in reinforcement learning for language models.

Key Characteristics

  • Base Model: Qwen3-1.7B, a 1.7 billion parameter model.
  • RL Algorithm: Trained using the GRPO (Generalized Reinforcement Policy Optimization) algorithm.
  • Reward Mechanism: Utilizes a frozen, prompt-specific initial rubric (R0) for training reward.
  • Domain: Specifically trained within the RaR Medicine domain, covering ten planned audit points through step 48 of the experiment.
  • Export Format: Provided in Hugging Face Transformers format, using BF16 safetensors.
  • Contents: Includes model weights, configuration, tokenizer, and a chat template.

Intended Use

This checkpoint is strictly a research artifact. Its primary purpose is for academic study into the effects of static rubrics and reward saturation during policy optimization processes. It is explicitly stated that these medicine checkpoints are not medical devices and should not be used as a substitute for professional medical advice. Developers should consider this model for research purposes related to RL fine-tuning and policy analysis, particularly in domain-specific contexts like medicine, rather than for general-purpose applications.