laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_warmup-ratio_0-05_Qwen3-32B
The laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_warmup-ratio_0-05_Qwen3-32B model is a 32 billion parameter language model fine-tuned from Qwen/Qwen3-32B. It was trained on the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset, suggesting a specialization in content related to Stack Exchange and Overflow sandboxes. With a context length of 32768 tokens, this model is likely optimized for tasks requiring deep understanding and generation within technical Q&A domains.
Loading preview...
Model Overview
This model, laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_warmup-ratio_0-05_Qwen3-32B, is a specialized large language model built upon the Qwen/Qwen3-32B architecture. It features 32 billion parameters and supports a substantial context length of 32768 tokens.
Training Focus
The model has undergone fine-tuning using the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset. This specific dataset choice indicates an optimization for tasks related to technical question-and-answer platforms, such as Stack Exchange and Overflow sandboxes, likely enhancing its ability to understand and generate relevant content in these domains.
Training Configuration
Key training hyperparameters include:
- Learning Rate: 4e-05
- Optimizer: AdamW_Torch_Fused
- Scheduler: Cosine with a 0.05 warmup ratio
- Epochs: 7.0
- Total Batch Size: 32 (across 16 GPUs with 2 gradient accumulation steps)
This configuration suggests a robust training process aimed at leveraging the base Qwen3-32B model's capabilities for specialized technical reasoning and content generation.