yuhan-nlp/verl-grpo-medium-qwen3-4b-step100-repro

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 13, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The yuhan-nlp/verl-grpo-medium-qwen3-4b-step100-repro is a 4 billion parameter Qwen3-based causal language model, fine-tuned using the GRPO method on the CollabLLM medium document-writing task. This model is an independent reproduction of a previous checkpoint, trained with the same recipe and data but on different hardware. It is specifically designed for document-writing tasks within the CollabLLM framework, offering a specialized solution for collaborative text generation.

Loading preview...

Model Overview

The yuhan-nlp/verl-grpo-medium-qwen3-4b-step100-repro is a 4 billion parameter language model based on the Qwen3 architecture. It has been fine-tuned using the GRPO (Generative Reinforcement Learning with Policy Optimization) method, specifically for the CollabLLM medium document-writing task. This particular model is a reproduction of an existing checkpoint, ensuring the same training recipe and data were used, but on an independent hardware setup.

Key Characteristics

  • Architecture: Based on Qwen/Qwen3-4B.
  • Training Method: Utilizes GRPO for fine-tuning, without an initial Supervised Fine-Tuning (SFT) warm start.
  • Task Specialization: Optimized for document-writing within the CollabLLM framework.
  • Reproducibility: Represents an independent reproduction, ensuring consistency with its original counterpart.
  • Direct Loading: Provided as a merged Hugging Face checkpoint, allowing for direct loading without an adapter step.

Training Data and Evaluation

The model was trained using the yuhan-nlp/collabllm-medium-rl-grpo dataset. While no specific metrics are reported directly with this model, detailed evaluation results, including per-example traces and judge outputs, are available in yuhan-nlp/collabllm-medium-outputs under benchmark_runs/. The evaluation setup involved using this model as an assistant, with Qwen/Qwen3.5-9B as the user simulator and Qwen/Qwen3.5-27B as the judge, across an evaluation size of 100 examples.