yuhan-nlp/verl-grpo-medium-qwen3-4b-step100-repro
The yuhan-nlp/verl-grpo-medium-qwen3-4b-step100-repro is a 4 billion parameter Qwen3-based causal language model, fine-tuned using the GRPO method on the CollabLLM medium document-writing task. This model is an independent reproduction of a previous checkpoint, trained with the same recipe and data but on different hardware. It is specifically designed for document-writing tasks within the CollabLLM framework, offering a specialized solution for collaborative text generation.
Loading preview...
Model Overview
The yuhan-nlp/verl-grpo-medium-qwen3-4b-step100-repro is a 4 billion parameter language model based on the Qwen3 architecture. It has been fine-tuned using the GRPO (Generative Reinforcement Learning with Policy Optimization) method, specifically for the CollabLLM medium document-writing task. This particular model is a reproduction of an existing checkpoint, ensuring the same training recipe and data were used, but on an independent hardware setup.
Key Characteristics
- Architecture: Based on
Qwen/Qwen3-4B. - Training Method: Utilizes GRPO for fine-tuning, without an initial Supervised Fine-Tuning (SFT) warm start.
- Task Specialization: Optimized for document-writing within the CollabLLM framework.
- Reproducibility: Represents an independent reproduction, ensuring consistency with its original counterpart.
- Direct Loading: Provided as a merged Hugging Face checkpoint, allowing for direct loading without an adapter step.
Training Data and Evaluation
The model was trained using the yuhan-nlp/collabllm-medium-rl-grpo dataset. While no specific metrics are reported directly with this model, detailed evaluation results, including per-example traces and judge outputs, are available in yuhan-nlp/collabllm-medium-outputs under benchmark_runs/. The evaluation setup involved using this model as an assistant, with Qwen/Qwen3.5-9B as the user simulator and Qwen/Qwen3.5-27B as the judge, across an evaluation size of 100 examples.