willamazon1/qwen3-8b-tmax-aenv-v39b-iter099
willamazon1/qwen3-8b-tmax-aenv-v39b-iter099 is an 8 billion parameter Qwen3-based language model, developed by willamazon1, specifically trained using reinforcement learning with group-relative policy optimization on asynchronous multi-turn agentic-environment rollouts. This checkpoint, taken at iteration 99, is fine-tuned from willamazon1/qwen3-8b-tmax-sft-v3-iter353. It is designed for agentic tasks and environments, leveraging its RL training for improved performance in interactive scenarios.
Loading preview...
Model Overview
willamazon1/qwen3-8b-tmax-aenv-v39b-iter099 is an 8 billion parameter Qwen3-based model, representing the 99th iteration of a reinforcement learning (RL) training run. It was initialized from willamazon1/qwen3-8b-tmax-sft-v3-iter353, which itself is a supervised fine-tuned (SFT) version of Qwen/Qwen3-8B.
Key Training Details
- Training Method: Reinforcement learning using group-relative policy optimization.
- Environment: Asynchronous multi-turn agentic-environment rollouts.
- Precision: Trained and provided in
bfloat16precision. - Architecture: Qwen3 with 36 layers, a hidden size of 4096, 32 attention heads, 8 KV heads, and a vocabulary size of 151936.
Unique Characteristics
This model is part of a series of checkpoints (iterations 69–219) from the tmax_aenv_v39b run, allowing for comparison of performance across different stages of RL training. The conversion process from Megatron-LM torch_dist to HuggingFace safetensors ensured data integrity, with checks for NaN/Inf values and successful loading and sampling prior to upload.
Intended Use Cases
Given its RL training on agentic environments, this model is particularly suited for applications requiring:
- Agentic tasks: Scenarios where the model needs to interact with an environment and make decisions.
- Multi-turn interactions: Handling complex, sequential dialogues or tasks.
- Comparative analysis: Researchers can utilize this specific iteration to study the effects of RL training progression by comparing it with other checkpoints from the same run.