willamazon1/qwen3-8b-tmax-aenv-v39b-iter099

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 18, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

willamazon1/qwen3-8b-tmax-aenv-v39b-iter099 is an 8 billion parameter Qwen3-based language model, developed by willamazon1, specifically trained using reinforcement learning with group-relative policy optimization on asynchronous multi-turn agentic-environment rollouts. This checkpoint, taken at iteration 99, is fine-tuned from willamazon1/qwen3-8b-tmax-sft-v3-iter353. It is designed for agentic tasks and environments, leveraging its RL training for improved performance in interactive scenarios.

Loading preview...

Model Overview

willamazon1/qwen3-8b-tmax-aenv-v39b-iter099 is an 8 billion parameter Qwen3-based model, representing the 99th iteration of a reinforcement learning (RL) training run. It was initialized from willamazon1/qwen3-8b-tmax-sft-v3-iter353, which itself is a supervised fine-tuned (SFT) version of Qwen/Qwen3-8B.

Key Training Details

  • Training Method: Reinforcement learning using group-relative policy optimization.
  • Environment: Asynchronous multi-turn agentic-environment rollouts.
  • Precision: Trained and provided in bfloat16 precision.
  • Architecture: Qwen3 with 36 layers, a hidden size of 4096, 32 attention heads, 8 KV heads, and a vocabulary size of 151936.

Unique Characteristics

This model is part of a series of checkpoints (iterations 69–219) from the tmax_aenv_v39b run, allowing for comparison of performance across different stages of RL training. The conversion process from Megatron-LM torch_dist to HuggingFace safetensors ensured data integrity, with checks for NaN/Inf values and successful loading and sampling prior to upload.

Intended Use Cases

Given its RL training on agentic environments, this model is particularly suited for applications requiring:

  • Agentic tasks: Scenarios where the model needs to interact with an environment and make decisions.
  • Multi-turn interactions: Handling complex, sequential dialogues or tasks.
  • Comparative analysis: Researchers can utilize this specific iteration to study the effects of RL training progression by comparing it with other checkpoints from the same run.