yuhan-nlp/verl-grpo-medium-qwen3-4b-step129-repro

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 13, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The yuhan-nlp/verl-grpo-medium-qwen3-4b-step129-repro model is a 4 billion parameter Qwen3-based causal language model, reproduced by yuhan-nlp, specifically trained with GRPO on the CollabLLM medium document-writing task. This model is designed for collaborative document generation, leveraging a unique training methodology without an SFT warm start. It is a direct merge of an 8-shard FSDP actor checkpoint, optimized for its specific document-writing application.

Loading preview...

Model Overview

This model, yuhan-nlp/verl-grpo-medium-qwen3-4b-step129-repro, is a reproduction of a 4 billion parameter Qwen3 model, developed by yuhan-nlp. It was trained using the GRPO (Gradient Regularized Policy Optimization) method on the CollabLLM medium document-writing task. A key characteristic is its training approach, which does not utilize a Supervised Fine-Tuning (SFT) warm start, distinguishing it from many other language models.

Key Capabilities

  • Collaborative Document Writing: Specifically trained for the CollabLLM medium document-writing task, indicating its specialization in generating coherent and contextually relevant text for collaborative environments.
  • GRPO Training: Leverages the GRPO training methodology, suggesting an optimization for specific policy-based learning objectives in language generation.
  • Direct Checkpoint: Provided as a merged Hugging Face checkpoint, allowing for direct loading without additional adapter steps.

Provenance and Evaluation

The model was merged from an 8-shard FSDP actor checkpoint (collabllm-qwen3-4B-medium-large-epoch1/global_step_129/actor) and trained on the yuhan-nlp/collabllm-medium-rl-grpo dataset. Evaluation metrics are not directly reported with the model; instead, detailed benchmark runs, including per-example traces and judge outputs, are available in yuhan-nlp/collabllm-medium-outputs under benchmark_runs/. The evaluation setup involved using this model as an assistant, with Qwen/Qwen3.5-9B as the user simulator and Qwen/Qwen3.5-27B as the judge.