HYU-NLP-EVAL/qwen3-1.7b-rar-science-static-r0-step-000
The HYU-NLP-EVAL/qwen3-1.7b-rar-science-static-r0-step-000 is a 1.7 billion parameter Qwen3-based causal language model, developed by HYU-NLP-EVAL, specifically fine-tuned using the GRPO algorithm for research into reward saturation and static-rubric staleness in the science domain. This model serves as a research artifact, providing a policy checkpoint from a static-rubric discriminability-horizon experiment. It is designed for studying policy optimization dynamics rather than general-purpose application, with a context length of 32768 tokens.
Loading preview...
Overview
This model, HYU-NLP-EVAL/qwen3-1.7b-rar-science-static-r0-step-000, is a 1.7 billion parameter Qwen3-based causal language model. It represents a specific policy checkpoint from a static-rubric discriminability-horizon experiment conducted by HYU-NLP-EVAL. The model was fine-tuned using the GRPO (Generalized Policy Optimization) algorithm, with training focused on a frozen prompt-specific initial rubric (R0) within the science domain.
Key Characteristics
- Base Model: Qwen/Qwen3-1.7B
- Fine-tuning Algorithm: GRPO (Generalized Policy Optimization)
- Training Reward: Frozen prompt-specific initial rubric (
R0) - Domain: Science, specifically for research into reward saturation and static-rubric staleness.
- Export Format: Hugging Face Transformers, BF16 safetensors, including model weights, configuration, tokenizer, and chat template.
- Context Length: 32768 tokens.
Intended Use
This model is primarily a research artifact designed for studying the dynamics of reward saturation and the staleness of static rubrics during policy optimization. It is not intended for general-purpose applications or as a production-ready model. The science training for this specific experiment was saved through step 3, with this repository containing a checkpoint from step 0.