tokyotech-llm/Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0010000
This model is a Llama-3.1-8B parameter language model from tokyotech-llm, continually pre-trained on 50 billion tokens with a specific focus on mathematical datasets from SwallowMath and multilingual text. It is designed to evaluate mathematical reasoning and problem-solving capabilities, incorporating a mix of mathematical, code, and multilingual data. The model is part of the SwallowMath ablation experiments, aiming to enhance performance in mathematical tasks.
Loading preview...
Model Overview
This model, developed by tokyotech-llm, is a continual pre-training of the Llama-3.1-8B architecture. It was trained on 50 billion tokens, specifically to evaluate and enhance performance in mathematical reasoning and problem-solving, as part of the SwallowMath ablation experiments (experiment 2).
Key Capabilities and Training
- Mathematical Reasoning: Optimized through a significant portion of mathematical datasets (4.8% SwallowMath, including rewritten data) within its training mix.
- Multilingual Support: Incorporates a large volume of multilingual text (82%) alongside code data (13.1%) to maintain broad language understanding.
- Architecture: Based on Llama-3.1, utilizing a Llama-3 tokenizer and trained with bfloat16 precision.
- Training Scale: Pre-trained on 50 billion tokens with a sequence length of 8,192, using 64 NVIDIA H100 GPUs.
Evaluation and Performance
The model was evaluated across a comprehensive suite of benchmarks, including mathematical reasoning (GSM8K, MATH), code generation (HumanEval), and general tasks (OpenBookQA, TriviaQA, HellaSwag, SQuAD 2.0, XWINO, MMLU, BBH). Performance metrics are provided for various checkpoints during the 50 billion token training, demonstrating its capabilities across these domains. For instance, at 50 billion tokens, it achieved 0.6535 on GSM8K and 0.3160 on MATH, indicating its specialized focus on mathematical tasks.