xuzishan/envace2.0-non-conv-rl-20260514-ckpt200
The xuzishan/envace2.0-non-conv-rl-20260514-ckpt200 model is an 8 billion parameter Qwen3ForCausalLM architecture developed by xuzishan, specifically a merged Hugging Face inference checkpoint from an EnvScaler non-conversation Reinforcement Learning (RL) run. This model is designed for inference and evaluation, representing a specific training step (200) of an RL process. It is suitable for applications requiring a Qwen3-8B class model fine-tuned through non-conversational reinforcement learning.
Loading preview...
EnvACE 2.0 Non-Conversation RL Model
This repository hosts the xuzishan/envace2.0-non-conv-rl-20260514-ckpt200 model, an 8 billion parameter language model based on the Qwen3ForCausalLM architecture. It represents a merged Hugging Face inference checkpoint derived from an EnvScaler non-conversation Reinforcement Learning (RL) run, specifically at training step 200.
Key Characteristics
- Architecture: Qwen3ForCausalLM (Qwen3-8B class).
- Origin: Produced from an EnvScaler non-conversation RL training process.
- Checkpoint: Represents training step 200 of the RL run.
- Contents: Includes four
safetensorsmodel shards, model index, configuration, tokenizer, generation configuration, and chat template. - Purpose: Designed for inference and evaluation tasks.
Important Notes
- This repository preserves the deployable merged model for direct use.
- It does not include the original distributed training checkpoint, optimizer state, scheduler state, or dataloader state.
- Exact training resume from step 200 is not possible with this repository's contents.