laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_adam-beta1_0-93_Qwen3-32B

TEXT GENERATIONPricing:Input $0.408 / Cached $0.0816 / Output $1.972Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jan 29, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

This model is a 32 billion parameter language model, fine-tuned from Qwen/Qwen3-32B by laion. It is specifically optimized for reasoning tasks, having been trained on the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset. With a context length of 32768 tokens, it is designed to handle complex problem-solving and logical inference scenarios.

Loading preview...

Model Overview

This model, developed by laion, is a fine-tuned version of the Qwen3-32B base model. It features 32 billion parameters and supports a context length of 32768 tokens. The fine-tuning process specifically utilized the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset.

Key Characteristics

  • Base Model: Qwen/Qwen3-32B
  • Parameter Count: 32 billion
  • Context Length: 32768 tokens
  • Training Data Focus: Reasoning tasks, leveraging the stackexchange-overflow-sandboxes dataset.

Training Details

The model was trained with a learning rate of 4e-05, a batch size of 32 (total), and utilized the ADAMW_TORCH_FUSED optimizer with specific beta parameters (0.93, 0.999). The training spanned 7 epochs with a cosine learning rate scheduler and a warmup ratio of 0.1. This configuration suggests an emphasis on robust and stable training for complex reasoning capabilities.

Potential Use Cases

  • Complex Reasoning: Ideal for applications requiring logical inference and problem-solving.
  • Question Answering: Particularly suited for detailed and nuanced responses based on provided context.
  • Technical Content Generation: Given its training data, it may perform well in generating or understanding technical explanations and solutions.