tokyotech-llm/Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0005000
tokyotech-llm/Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0005000 is an 8 billion parameter Llama-3.1 model continually pre-trained by tokyotech-llm. This model focuses on mathematical reasoning and problem-solving, having been trained on 50 billion tokens including a significant mix of mathematical datasets from SwallowMath and multilingual text. It is specifically designed to evaluate performance in mathematical tasks as part of the SwallowMath ablation experiments, offering enhanced capabilities in this domain.
Loading preview...
Model Overview
This model, developed by tokyotech-llm, is a continually pre-trained version of the Llama-3.1-8B architecture. It was trained on 50 billion tokens, incorporating a specialized mix of 4.8% mathematical datasets from SwallowMath (finemath-4+ rewritten), 13.1% code, and 82% multilingual text. This specific training configuration is part of the SwallowMath ablation experiments (experiment 2), aimed at evaluating and enhancing mathematical reasoning and problem-solving capabilities.
Key Capabilities
- Enhanced Mathematical Reasoning: Optimized through targeted pre-training on mathematical datasets, making it suitable for complex math problems.
- Code Generation: Includes a substantial portion of code data in its training mix, contributing to its ability in code-related tasks.
- Multilingual Understanding: Benefits from a broad multilingual text corpus, supporting diverse language applications.
- Llama-3.1 Foundation: Built upon the robust Llama-3.1 architecture, providing a strong base for various NLP tasks.
Training Details
The model was trained using Megatron-LM on 64 NVIDIA H100 GPUs. Its training data included 2.4 billion mathematical tokens, 6.5 billion code tokens (SwallowCode), and 40.66 billion multilingual text tokens (Japanese Wikipedia, Swallow Corpus v2, Laboro-ParaCorpus, English Wikipedia, Cosmopedia, DCLM). Evaluation results are available across various benchmarks, including GSM8K, MATH, HumanEval, and general tasks like MMLU and BBH, with performance metrics reported at different token checkpoints up to 50 billion tokens.