xuzishan/envace2.0-non-conv-rl-grpo-32gpu-20260520-ckpt200

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 12, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The xuzishan/envace2.0-non-conv-rl-grpo-32gpu-20260520-ckpt200 is an 8 billion parameter Qwen3ForCausalLM model checkpoint, developed by xuzishan. This model is a merged Hugging Face inference checkpoint from an aligned EnvScaler non-conversation GRPO training run. It is designed for inference and evaluation tasks, providing a deployable model for non-conversational applications.

Loading preview...

EnvACE 2.0 Non-Conversation GRPO Model

This repository hosts the xuzishan/envace2.0-non-conv-rl-grpo-32gpu-20260520-ckpt200 model, an 8 billion parameter checkpoint based on the Qwen3ForCausalLM architecture. It represents a merged Hugging Face inference checkpoint derived from an aligned EnvScaler non-conversation GRPO (Generalized Reinforcement Learning Policy Optimization) training run, specifically at training step 200.

Key Characteristics

  • Architecture: Utilizes the Qwen3ForCausalLM class, indicating a Qwen3-8B base.
  • Format: Comprises four safetensors model shards, along with necessary configuration files including model index, tokenizer, generation configuration, and chat template.
  • Origin: This checkpoint is a direct export from the envscaler_non_conv_rl_grpo_32gpu_aligned_20260520_055146 run.
  • Integrity: All 14 model and runtime files have been byte-for-byte verified against their local counterparts, ensuring data integrity.

Use Cases and Limitations

This model is primarily intended for inference and evaluation. It provides a deployable merged model suitable for running predictions and assessing performance in non-conversational contexts. It is important to note that this repository does not include the original distributed training checkpoint, optimizer state, scheduler state, or dataloader state. Therefore, it cannot be used to reproduce an exact training resume from step 200, but it is fully functional for its intended inference and evaluation purposes.