laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_warmup-ratio_0-01_Qwen3-32B

TEXT GENERATIONPricing:Input $0.408 / Cached $0.0816 / Output $1.972Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jan 20, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_warmup-ratio_0-01_Qwen3-32B model is a 32 billion parameter language model fine-tuned from Qwen/Qwen3-32B. It was trained on the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset, suggesting a specialization in content related to StackExchange and Overflow sandboxes. With a context length of 32768 tokens, this model is likely optimized for tasks requiring deep understanding and generation within technical Q&A domains.

Loading preview...

Model Overview

This model, laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_warmup-ratio_0-01_Qwen3-32B, is a specialized fine-tuned version of the Qwen/Qwen3-32B base model. It leverages a 32 billion parameter architecture and supports a substantial context window of 32,768 tokens.

Training Focus

The model was fine-tuned using the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset. This specific dataset choice indicates an optimization for tasks and content typically found on platforms like StackExchange and Overflow, suggesting a strong capability in technical question-answering, code-related discussions, and problem-solving within sandbox environments.

Training Configuration

Key training hyperparameters include:

  • Learning Rate: 4e-05
  • Batch Size: 1 (train), 8 (eval) with 2 gradient accumulation steps, leading to a total train batch size of 32.
  • Optimizer: AdamW_Torch_Fused
  • Scheduler: Cosine learning rate scheduler with a 0.01 warmup ratio.
  • Epochs: 7.0

This configuration aims to refine the base Qwen3-32B model's understanding and generation capabilities for highly specific technical and reasoning-intensive dialogues.