xuzishan/envace2.0-non-conv-rl-20260514-ckpt200

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 12, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The xuzishan/envace2.0-non-conv-rl-20260514-ckpt200 model is an 8 billion parameter Qwen3ForCausalLM architecture developed by xuzishan, specifically a merged Hugging Face inference checkpoint from an EnvScaler non-conversation Reinforcement Learning (RL) run. This model is designed for inference and evaluation, representing a specific training step (200) of an RL process. It is suitable for applications requiring a Qwen3-8B class model fine-tuned through non-conversational reinforcement learning.

Loading preview...

EnvACE 2.0 Non-Conversation RL Model

This repository hosts the xuzishan/envace2.0-non-conv-rl-20260514-ckpt200 model, an 8 billion parameter language model based on the Qwen3ForCausalLM architecture. It represents a merged Hugging Face inference checkpoint derived from an EnvScaler non-conversation Reinforcement Learning (RL) run, specifically at training step 200.

Key Characteristics

  • Architecture: Qwen3ForCausalLM (Qwen3-8B class).
  • Origin: Produced from an EnvScaler non-conversation RL training process.
  • Checkpoint: Represents training step 200 of the RL run.
  • Contents: Includes four safetensors model shards, model index, configuration, tokenizer, generation configuration, and chat template.
  • Purpose: Designed for inference and evaluation tasks.

Important Notes

  • This repository preserves the deployable merged model for direct use.
  • It does not include the original distributed training checkpoint, optimizer state, scheduler state, or dataloader state.
  • Exact training resume from step 200 is not possible with this repository's contents.