tokyotech-llm/Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0007500
The tokyotech-llm/Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0007500 is an 8 billion parameter Llama-3.1 model continually pre-trained on 50 billion tokens, including a significant mix of mathematical and code datasets, with a 32768 token context length. Developed by tokyotech-llm, this model is specifically designed to evaluate and enhance mathematical reasoning and problem-solving capabilities. It integrates 4.8% mathematical data and 13.1% code data into its training mix, making it suitable for tasks requiring strong analytical and computational skills.
Loading preview...
Model Overview
This model, tokyotech-llm/Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0007500, is a continually pre-trained Llama-3.1-8B variant developed by tokyotech-llm. It was trained on 50 billion tokens with a focus on mathematical reasoning and problem-solving, as part of the SwallowMath ablation experiments.
Key Characteristics
- Base Model: Llama-3.1-8B architecture.
- Training Data Mix: Includes 4.8% mathematical datasets (from SwallowMath), 13.1% code data, and 82% multilingual text.
- Purpose: Designed to evaluate and improve performance in mathematical reasoning and problem-solving, following the methodology outlined in the SwallowMath paper.
- Context Length: Supports a sequence length of 8,192 tokens during training.
- Evaluation: Benchmarked across various tasks including mathematical reasoning (GSM8K, MATH), code generation (HumanEval), and general tasks (MMLU, BBH, etc.).
Performance Highlights
Evaluation results at 50 billion tokens show:
- GSM8K: 0.6535
- MATH: 0.3160
- HumanEval: 0.3683
- MMLU: 0.6337
Use Cases
This model is particularly well-suited for:
- Mathematical Problem Solving: Excels in tasks requiring numerical reasoning and equation solving.
- Code Generation: Benefits from a substantial code dataset in its training mix.
- Research in LLM Capabilities: Ideal for researchers exploring the impact of specialized data mixes on mathematical and logical reasoning in large language models.