laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_adam-beta1_0-91_Qwen3-32B

TEXT GENERATIONPricing:Input $0.408 / Cached $0.0816 / Output $1.972Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jan 29, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_adam-beta1_0-91_Qwen3-32B is a 32 billion parameter language model, fine-tuned from Qwen/Qwen3-32B. It was specifically trained on the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset, indicating an optimization for reasoning tasks, particularly those found in StackExchange and Overflow sandboxes. This model is designed to enhance performance in complex problem-solving and knowledge-based question answering within technical domains.

Loading preview...

Model Overview

This model, laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_adam-beta1_0-91_Qwen3-32B, is a 32 billion parameter language model. It is a fine-tuned variant of the Qwen/Qwen3-32B base model, developed by laion. The fine-tuning process utilized the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset, suggesting a specialization in reasoning and problem-solving tasks, likely within technical Q&A contexts such as StackExchange and Overflow.

Training Details

The model was trained with specific hyperparameters to optimize its performance:

  • Learning Rate: 4e-05
  • Optimizer: ADAMW_TORCH_FUSED with betas=(0.91, 0.999)
  • Batch Size: A total training batch size of 32 (with gradient accumulation steps of 2)
  • Epochs: 7.0
  • Scheduler: Cosine learning rate scheduler with a warmup ratio of 0.1

Key Characteristics

  • Base Model: Qwen/Qwen3-32B (32 billion parameters)
  • Specialization: Fine-tuned on a dataset focused on reasoning, likely improving its ability to handle complex queries and provide detailed explanations, particularly in technical or problem-solving scenarios.

Potential Use Cases

Given its fine-tuning dataset, this model is likely well-suited for:

  • Technical Q&A: Answering questions and providing solutions similar to those found on StackExchange or Overflow.
  • Reasoning Tasks: Engaging in complex logical reasoning and problem-solving.
  • Knowledge Retrieval: Extracting and synthesizing information from technical documentation or discussions.