tokyotech-llm/Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0007500
The tokyotech-llm/Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0007500 is an 8 billion parameter Llama-3.1 model continually pre-trained by tokyotech-llm. It was trained on 50 billion tokens, including a significant mix of mathematical datasets (Finemath-4+), code, and multilingual text, with a context length of 32768 tokens. This model is specifically designed to evaluate and enhance mathematical reasoning and problem-solving capabilities, making it suitable for tasks requiring strong analytical skills.
Loading preview...
Model Overview
This model, tokyotech-llm/Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0007500, is an 8 billion parameter Llama-3.1 variant continually pre-trained by tokyotech-llm. It is part of the SwallowMath ablation experiments, specifically designed to evaluate and improve performance in mathematical reasoning and problem-solving.
Key Characteristics
- Architecture: Based on Llama-3.1, with a Llama-3 tokenizer.
- Training Data: Continually pre-trained on 50 billion tokens, comprising a unique mix:
- ~4.8% mathematical data (Finemath-4+)
- ~13.1% code data (SwallowCode)
- ~82% multilingual text (Japanese and English corpora).
- Context Length: Supports a sequence length of 8,192 tokens.
- Purpose: Developed to assess and enhance mathematical reasoning and problem-solving, as detailed in the SwallowMath paper.
Evaluation & Performance
The model was evaluated across various benchmarks, including mathematical reasoning (GSM8K, MATH), code generation (HumanEval), and general tasks (MMLU, BBH, OpenBookQA, TriviaQA, HellaSwag, SQuAD 2.0, XWINO). Performance metrics are available for checkpoints up to 50 billion tokens, showing progressive improvement in mathematical and general reasoning tasks.
Use Cases
This model is particularly well-suited for applications requiring:
- Mathematical Problem Solving: Excels in tasks involving arithmetic, algebra, and complex mathematical reasoning.
- Code-Related Tasks: Benefits from its significant code training data, making it useful for code understanding or generation.
- Research in LLM Capabilities: Ideal for researchers exploring the impact of data mix on mathematical and general reasoning abilities in large language models.