xuzishan/envscaler-qwen3-8b-grpo-step63-20260515

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 6, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The xuzishan/envscaler-qwen3-8b-grpo-step63-20260515 is an 8.19 billion parameter Qwen3 causal language model, derived from an early EnvScaler non-conversational GRPO training run. This model is a recoverable checkpoint from step 63 of a 32-GPU training experiment, not selected based on formal benchmarks. It is suitable for research into early-stage reinforcement learning from human feedback (RLHF) models and their training dynamics.

Loading preview...

EnvScaler Qwen3-8B GRPO Step 63

This model, xuzishan/envscaler-qwen3-8b-grpo-step63-20260515, is a Hugging Face-format export of an early checkpoint from an EnvScaler non-conversational 32-GPU GRPO (Generative Reinforcement Learning with Policy Optimization) training run. It is based on the Qwen3 causal language model architecture, featuring approximately 8.19 billion parameters.

Key Characteristics

  • Provenance: Represents step 63 of an internal training run (envscaler_non_conv_rl_grpo_32gpu_20260515_225741).
  • Training: Trained between 2026-05-15 and 2026-05-16, completing 64 steps.
  • Non-Benchmark Selected: This checkpoint was not chosen based on formal evaluation workflows like BFCL, VitaBench, or Tau2, and should not be interpreted as a benchmark-optimized release.
  • Training Metrics: Achieved a final rollout score of 0.660812 at step 63, with a peak of 0.851259 at step 41. These are online training signals, not held-out benchmark results.
  • Format: Converted from a two-way tensor-parallel Megatron checkpoint to Hugging Face safetensors, with all 399 expected Qwen3 parameters validated.

Use Cases

This model is primarily suitable for:

  • Research: Investigating the behavior and performance of early-stage GRPO-trained models.
  • Training Analysis: Studying the progression of online training signals in reinforcement learning setups.
  • Experimental Development: As a base for further fine-tuning or experimentation where a non-benchmark-selected checkpoint is acceptable.