laion/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning_adam-beta1_0-93_Qwen3-32B
This model is a 32 billion parameter language model, fine-tuned from Qwen/Qwen3-32B by laion. It is specifically optimized for reasoning tasks, having been trained on the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset. With a context length of 32768 tokens, it is designed to handle complex problem-solving and logical inference scenarios.
Loading preview...
Model Overview
This model, developed by laion, is a fine-tuned version of the Qwen3-32B base model. It features 32 billion parameters and supports a context length of 32768 tokens. The fine-tuning process specifically utilized the open-athena/GLM-4.6-stackexchange-overflow-sandboxes-32eps-65k-reasoning dataset.
Key Characteristics
- Base Model: Qwen/Qwen3-32B
- Parameter Count: 32 billion
- Context Length: 32768 tokens
- Training Data Focus: Reasoning tasks, leveraging the
stackexchange-overflow-sandboxesdataset.
Training Details
The model was trained with a learning rate of 4e-05, a batch size of 32 (total), and utilized the ADAMW_TORCH_FUSED optimizer with specific beta parameters (0.93, 0.999). The training spanned 7 epochs with a cosine learning rate scheduler and a warmup ratio of 0.1. This configuration suggests an emphasis on robust and stable training for complex reasoning capabilities.
Potential Use Cases
- Complex Reasoning: Ideal for applications requiring logical inference and problem-solving.
- Question Answering: Particularly suited for detailed and nuanced responses based on provided context.
- Technical Content Generation: Given its training data, it may perform well in generating or understanding technical explanations and solutions.