xuzishan/envace2.0-non-conv-rl-grpo-32gpu-20260520-ckpt200
The xuzishan/envace2.0-non-conv-rl-grpo-32gpu-20260520-ckpt200 is an 8 billion parameter Qwen3ForCausalLM model checkpoint, developed by xuzishan. This model is a merged Hugging Face inference checkpoint from an aligned EnvScaler non-conversation GRPO training run. It is designed for inference and evaluation tasks, providing a deployable model for non-conversational applications.
Loading preview...
EnvACE 2.0 Non-Conversation GRPO Model
This repository hosts the xuzishan/envace2.0-non-conv-rl-grpo-32gpu-20260520-ckpt200 model, an 8 billion parameter checkpoint based on the Qwen3ForCausalLM architecture. It represents a merged Hugging Face inference checkpoint derived from an aligned EnvScaler non-conversation GRPO (Generalized Reinforcement Learning Policy Optimization) training run, specifically at training step 200.
Key Characteristics
- Architecture: Utilizes the
Qwen3ForCausalLMclass, indicating a Qwen3-8B base. - Format: Comprises four
safetensorsmodel shards, along with necessary configuration files including model index, tokenizer, generation configuration, and chat template. - Origin: This checkpoint is a direct export from the
envscaler_non_conv_rl_grpo_32gpu_aligned_20260520_055146run. - Integrity: All 14 model and runtime files have been byte-for-byte verified against their local counterparts, ensuring data integrity.
Use Cases and Limitations
This model is primarily intended for inference and evaluation. It provides a deployable merged model suitable for running predictions and assessing performance in non-conversational contexts. It is important to note that this repository does not include the original distributed training checkpoint, optimizer state, scheduler state, or dataloader state. Therefore, it cannot be used to reproduce an exact training resume from step 200, but it is fully functional for its intended inference and evaluation purposes.