willamazon1/qwen3-8b-tmax-aenv-v39b-iter209

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 18, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

willamazon1/qwen3-8b-tmax-aenv-v39b-iter209 is an 8 billion parameter Qwen3-based causal language model developed by willamazon1. This model is a reinforcement learning checkpoint, specifically iteration 209, from the tmax_aenv_v39b run. It was trained using group-relative policy optimization on asynchronous multi-turn agentic-environment rollouts, making it suitable for agentic applications and environments.

Loading preview...

Model Overview

This model, willamazon1/qwen3-8b-tmax-aenv-v39b-iter209, is an 8 billion parameter Qwen3-based causal language model. It represents a reinforcement learning (RL) checkpoint, specifically iteration 209, from the tmax_aenv_v39b training run. The model was initialized from willamazon1/qwen3-8b-tmax-sft-v3-iter353, which itself was a supervised fine-tuned (SFT) version of Qwen/Qwen3-8B.

Key Characteristics

  • Architecture: Qwen3 with 36 layers, a hidden size of 4096, 32 attention heads, and 8 KV heads. The vocabulary size is 151936.
  • Training Stage: Reinforcement Learning, utilizing an agentic environment with asynchronous multi-turn rollouts and group-relative policy optimization.
  • Precision: Trained and provided in bfloat16 precision.
  • Context Length: Supports a context length of 32768 tokens.

Intended Use Cases

This model is particularly well-suited for applications requiring agentic behavior and interaction within simulated environments, given its RL training methodology. Developers can leverage this checkpoint to explore and compare performance across different training iterations, as other checkpoints from iterations 69–219 of this run are also available.