seomh/Qwen3-8B-Base-OpenThoughts3Math-SFT-step750-lr5e-6
The seomh/Qwen3-8B-Base-OpenThoughts3Math-SFT-step750-lr5e-6 is an 8 billion parameter Qwen3-Base model fine-tuned for mathematical reasoning. This intermediate checkpoint, trained on the OpenThoughts3 math dataset, specializes in reproducing detailed thought processes ending in a boxed answer. It utilizes a plain ChatML template to preserve reasoning traces, making it distinct from standard Qwen3 models that might strip intermediate thoughts. This model is optimized for tasks requiring explicit step-by-step mathematical problem-solving.
Loading preview...
Model Overview
This model, seomh/Qwen3-8B-Base-OpenThoughts3Math-SFT-step750-lr5e-6, is an 8 billion parameter variant of the Qwen3-Base architecture. It is an intermediate checkpoint (step 750 of 2000) specifically fine-tuned for mathematical reasoning using the knowledge-distillation/openthoughts3_math dataset, which comprises 103,760 two-turn conversations.
Key Capabilities & Training Details
- Mathematical Reasoning: The model is trained to reproduce detailed thought processes, where every assistant message includes a
<think>...</think>trace culminating in a\boxed{}answer. - Custom ChatML Template: Unlike standard Qwen3 templates, this model uses a plain ChatML template that preserves all
<think>tags, ensuring the reasoning steps are maintained in the training targets. - Training Configuration:
- Base Model:
Qwen/Qwen3-8B-Base - Dataset:
knowledge-distillation/openthoughts3_math - Sequence Length: 32768 tokens (mean 13.8k, p99 16.8k, max 16.8k tokens in data)
- Optimizer: AdamW with a learning rate of 5e-6, cosine schedule, 3% warmup, 0.01 weight decay, and 1.0 gradient clipping.
- Precision: bf16, utilizing FSDP across 3 GPUs with verl.
- Base Model:
Important Notes for Usage
- The model's
generation_config.jsonandconfig.jsonspecify<|im_end|>(token 151645) as the turn terminator. The stockQwen3-8B-Base'sgeneration_config.jsonincorrectly uses 151643, which can cause non-termination issues with vLLM. Users should ensure the correct terminator token is used. - It is recommended to generate with
max_new_tokensaround 16384 to accommodate the typical output length.