laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_warmup-ratio_0-05_Qwen3-32B

TEXT GENERATIONPricing:Input $0.408 / Cached $0.0816 / Output $1.972Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jan 20, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_warmup-ratio_0-05_Qwen3-32B model is a 32 billion parameter language model fine-tuned from Qwen/Qwen3-32B. It was trained on the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset, suggesting a specialization in content related to Stack Exchange and Overflow sandboxes. With a context length of 32768 tokens, this model is likely optimized for tasks requiring deep understanding and generation within technical Q&A domains.

Loading preview...

Model Overview

This model, laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_warmup-ratio_0-05_Qwen3-32B, is a specialized large language model built upon the Qwen/Qwen3-32B architecture. It features 32 billion parameters and supports a substantial context length of 32768 tokens.

Training Focus

The model has undergone fine-tuning using the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset. This specific dataset choice indicates an optimization for tasks related to technical question-and-answer platforms, such as Stack Exchange and Overflow sandboxes, likely enhancing its ability to understand and generate relevant content in these domains.

Training Configuration

Key training hyperparameters include:

  • Learning Rate: 4e-05
  • Optimizer: AdamW_Torch_Fused
  • Scheduler: Cosine with a 0.05 warmup ratio
  • Epochs: 7.0
  • Total Batch Size: 32 (across 16 GPUs with 2 gradient accumulation steps)

This configuration suggests a robust training process aimed at leveraging the base Qwen3-32B model's capabilities for specialized technical reasoning and content generation.