tokyotech-llm/Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0012500
The tokyotech-llm/Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0012500 is an 8 billion parameter Llama-3.1 model continually pre-trained by tokyotech-llm. It is specifically optimized for mathematical reasoning and problem-solving, incorporating a mix of mathematical, code, and multilingual datasets. This model is part of the SwallowMath ablation experiments, focusing on evaluating performance in mathematical tasks. It features a 32768 token context length and was trained on 50 billion tokens.
Loading preview...
Model Overview
This model, developed by tokyotech-llm, is a continually pre-trained version of the Llama-3.1-8B architecture. It was specifically designed as part of the SwallowMath ablation experiments (experiment 2) to evaluate and enhance performance in mathematical reasoning and problem-solving tasks.
Key Characteristics
- Architecture: Based on Llama-3.1 with 8 billion parameters.
- Training Data: Continually pre-trained on 50 billion tokens, comprising a specialized mix:
- ~4.8% mathematical data from SwallowMath (finemath-4+ rewritten).
- ~13.1% code data (SwallowCode).
- ~82% multilingual text, including Japanese and English corpora.
- Context Length: Supports a sequence length of 8,192 tokens.
- Purpose: Primarily aimed at improving and assessing mathematical capabilities, as detailed in the SwallowMath paper.
Evaluation Highlights
The model's performance was evaluated across various benchmarks, including mathematical reasoning (GSM8K, MATH), code generation (HumanEval), and general tasks (MMLU, BBH, OpenBookQA, TriviaQA, HellaSwag, SQuAD 2.0, XWINO). The evaluation results, reported at different training checkpoints up to 50 billion tokens, demonstrate its progression in these areas, particularly in mathematical problem-solving.