willamazon1/Qwen3-8B-SDFT-Math-LoRA-iter45
willamazon1/Qwen3-8B-SDFT-Math-LoRA-iter45 is an 8.2 billion parameter causal language model from the Qwen3 series, pre-trained on 36 trillion tokens across 119 languages with a 32,768 token context length. This model incorporates advanced training techniques like qk layernorm and a three-stage pre-training process focusing on general knowledge, STEM/coding reasoning, and long-context comprehension. It is designed for broad language modeling and tasks requiring strong reasoning skills, particularly in STEM and coding domains.
Loading preview...
Model Overview
willamazon1/Qwen3-8B-SDFT-Math-LoRA-iter45 is an 8.2 billion parameter causal language model based on the Qwen3 architecture, developed by the Qwen Team. This model is part of the latest generation of Qwen series, building upon significant advancements in training data, model architecture, and optimization techniques. It features a substantial pre-training corpus of 36 trillion tokens covering 119 languages, a threefold increase in language coverage compared to Qwen2.5, with a rich mix of high-quality data including coding, STEM, reasoning, and multilingual content.
Key Features and Training
- Expanded Pre-training Corpus: Trained on 36 trillion tokens across 119 languages, emphasizing high-quality coding, STEM, reasoning, and multilingual data.
- Architectural Refinements: Incorporates advanced training techniques such as qk layernorm for improved stability and performance across all models.
- Three-stage Pre-training: A structured approach where Stage 1 focuses on general language modeling, Stage 2 enhances reasoning skills (STEM, coding, logical reasoning), and Stage 3 extends long-context comprehension up to 32,768 tokens.
- Optimized Hyperparameter Tuning: Utilizes scaling law studies to systematically tune hyperparameters for better training dynamics and performance.
- Context Length: Supports a substantial context window of 32,768 tokens.
Use Cases
This model is well-suited for applications requiring robust general language understanding, advanced reasoning capabilities, particularly in STEM and coding domains, and processing long textual inputs. Its extensive multilingual training also makes it suitable for diverse language applications.