HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-016

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 3, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-016 is a 2 billion parameter Qwen3-1.7B policy checkpoint, developed by HYU-NLP-EVAL, derived from the Qwen/Qwen3-1.7B base model. This model is a research artifact from an experiment studying reward saturation and static-rubric staleness during policy optimization, specifically within the RaR Medicine domain. It is fine-tuned using the GRPO RL algorithm with a frozen prompt-specific initial rubric (R0). This checkpoint is intended for research into policy optimization dynamics rather than direct application.

Loading preview...

Model Overview

HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-016 is a 2 billion parameter policy checkpoint based on the Qwen/Qwen3-1.7B model. It is a research artifact from an experiment conducted by HYU-NLP-EVAL, focusing on the dynamics of reward saturation and static-rubric staleness during policy optimization.

Key Characteristics

  • Base Model: Qwen/Qwen3-1.7B
  • RL Algorithm: Trained using GRPO (Generalized Reward Policy Optimization).
  • Reward Mechanism: Utilizes a frozen prompt-specific initial rubric (R0) for training.
  • Domain: Specifically trained within the RaR Medicine domain, with this checkpoint representing step 16 of the optimization process.
  • Export Format: Provided in Hugging Face Transformers format, BF16 safetensors, including model weights, configuration, tokenizer, and chat template.

Intended Use

This model is primarily intended for academic and research purposes, specifically for:

  • Studying reward saturation in reinforcement learning policies.
  • Investigating static-rubric staleness during policy optimization.

Important Note: Medicine checkpoints like this one are research artifacts and are not medical devices. They must not be used as a substitute for professional medical advice or in any clinical application.