tokyotech-llm/Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0010000
The tokyotech-llm/Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0010000 is an 8 billion parameter Llama-3.1 model continually pre-trained by tokyotech-llm. It was trained on 50 billion tokens, including a significant mix of mathematical, code, and multilingual text datasets, with a 32K context length. This model is specifically designed to evaluate and enhance mathematical reasoning and problem-solving capabilities, serving as part of the SwallowMath ablation experiments.
Loading preview...
Model Overview
This model, developed by tokyotech-llm, is a continually pre-trained version of the Llama-3.1-8B architecture. It was trained on 50 billion tokens, focusing on evaluating and improving mathematical reasoning and problem-solving skills as part of the SwallowMath ablation experiments.
Key Training Details
- Base Model: Llama-3.1-8B
- Total Pretraining Tokens: 50 billion
- Data Mix: Approximately 4.8% mathematical data (Finemath-4+), 13.1% code data (SwallowCode), and 82% multilingual text (Japanese and English corpora).
- Context Length: 8,192 tokens during training, with a reported 32K context length.
- Hardware: Trained on 64 NVIDIA H100 GPUs using the TSUBAME supercomputer.
Evaluation & Performance
The model was evaluated across a range of benchmarks, including mathematical reasoning (GSM8K, MATH), code generation (HumanEval), and general tasks (MMLU, BBH, OpenBookQA, TriviaQA, HellaSwag, SQuAD 2.0, XWINO). Performance metrics are provided for checkpoints up to 50 billion tokens, showing progressive improvements in mathematical and general reasoning tasks.
Use Cases
This model is particularly suitable for research and applications requiring:
- Mathematical Reasoning: Solving equations and complex math problems.
- Code Generation: Assisting with programming tasks.
- Multilingual Text Processing: Handling tasks involving both English and Japanese text.