yuhan-nlp/verl-grpo-medium-qwen3-4b-step50-repro

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 13, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The yuhan-nlp/verl-grpo-medium-qwen3-4b-step50-repro model is a 4 billion parameter Qwen3-based causal language model, a reproduction of a prior model trained with GRPO on the CollabLLM medium document-writing task. Developed by yuhan-nlp, this model is specifically designed for collaborative document generation, leveraging the verl CollabLLM recipe without an SFT warm start. It is optimized for tasks requiring structured text output in a collaborative writing context, offering a 32768 token context length.

Loading preview...

Model Overview

This model, yuhan-nlp/verl-grpo-medium-qwen3-4b-step50-repro, is a 4 billion parameter Qwen3-based causal language model. It represents an independent-cluster reproduction of yuhan-nlp/verl-grpo-medium-qwen3-4b-step50, utilizing the same training recipe and data but on different hardware. The model was trained using GRPO (Generative Reinforcement Learning with Policy Optimization) on the CollabLLM medium document-writing task, specifically employing the verl CollabLLM recipe without an initial Supervised Fine-Tuning (SFT) warm start.

Key Characteristics

  • Base Model: Qwen/Qwen3-4B architecture.
  • Training Method: GRPO for reinforcement learning.
  • Task Focus: Collaborative document writing, as defined by the CollabLLM medium task.
  • Data: Trained on the yuhan-nlp/collabllm-medium-rl-grpo dataset.
  • Context Length: Supports a context window of 32768 tokens.
  • Deployment: Provided as a merged Hugging Face checkpoint, allowing direct loading without an adapter step.

Usage Notes

The model's chat template is included as chat_template.jinja and is automatically picked up by AutoTokenizer in transformers versions 4.57 and above. Evaluation metrics are not directly reported within the model card; instead, detailed per-example traces, judge outputs, and summary JSONs are available in yuhan-nlp/collabllm-medium-outputs under benchmark_runs/.