laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_global-batch-size_64_Qwen3-32B

TEXT GENERATIONPricing:Input $0.408 / Cached $0.0816 / Output $1.972Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Dec 23, 2025License:otherArchitecture:Transformer Featherless Exclusive Cold

The laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_global-batch-size_64_Qwen3-32B model is a 32 billion parameter language model fine-tuned from Qwen/Qwen3-32B. It was trained on the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset, suggesting a specialization in processing and generating content related to Stack Exchange and similar Q&A platforms. This model is likely optimized for tasks requiring detailed technical explanations, problem-solving, and information retrieval within specific knowledge domains, leveraging its 32768 token context length.

Loading preview...

Model Overview

This model, GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_global-batch-size_64_Qwen3-32B, is a specialized large language model built upon the Qwen3-32B architecture. It has been fine-tuned using the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset, indicating a strong focus on content derived from Stack Exchange and similar technical Q&A forums.

Training Details

The model underwent training with specific hyperparameters:

  • Base Model: Qwen/Qwen3-32B
  • Dataset: open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning
  • Learning Rate: 4e-05
  • Optimizer: AdamW_Torch_Fused
  • Epochs: 5.0
  • Total Batch Size: 64 (with gradient accumulation steps of 4 across 16 devices)

Potential Use Cases

Given its training data, this model is likely well-suited for:

  • Generating detailed answers to technical questions.
  • Summarizing discussions from developer forums.
  • Assisting with code-related queries and explanations.
  • Information retrieval within specific technical domains.

Further details on intended uses, limitations, and evaluation data are not provided in the current model card.