HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-048

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 3, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-048 is a 1.7 billion parameter Qwen3-based language model, fine-tuned using the GRPO reinforcement learning algorithm. This specific checkpoint, from step 48, is optimized for research into reward saturation and static-rubric staleness within the medical domain. It serves as a research artifact for studying policy optimization dynamics rather than a general-purpose medical advice tool.

Loading preview...

Overview

This model, HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-048, is a 1.7 billion parameter policy checkpoint derived from the Qwen/Qwen3-1.7B base model. It was developed by HYU-NLP-EVAL as part of an experiment investigating reward saturation and static-rubric staleness in policy optimization using the GRPO (Generalized Reinforcement Policy Optimization) algorithm.

Key Characteristics

  • Base Model: Qwen3-1.7B, a causal language model.
  • Training Method: Fine-tuned using the GRPO reinforcement learning algorithm.
  • Reward Mechanism: Utilizes a frozen, prompt-specific initial rubric (R0).
  • Domain Focus: Specifically trained within the Medicine domain, covering ten planned audit points through step 48.
  • Format: Exported in Hugging Face Transformers format, using BF16 safetensors.
  • Contents: Includes model weights, configuration, tokenizer, and chat template.

Intended Use

This checkpoint is primarily a research artifact.

  • Study Focus: Designed for studying reward saturation and the staleness of static rubrics during policy optimization.
  • Limitations: It is explicitly stated that these medicine checkpoints are not medical devices and should not be used as a substitute for professional medical advice. Its purpose is purely for experimental analysis within the NLP research community.