tokyotech-llm/Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0010000

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 10, 2025License:llama3.3Architecture:Transformer Featherless Exclusive Cold

The tokyotech-llm/Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0010000 is an 8 billion parameter Llama-3.1 model continually pre-trained by tokyotech-llm. It was trained on 50 billion tokens, including a significant mix of mathematical, code, and multilingual text datasets, with a 32K context length. This model is specifically designed to evaluate and enhance mathematical reasoning and problem-solving capabilities, serving as part of the SwallowMath ablation experiments.

Loading preview...

Model Overview

This model, developed by tokyotech-llm, is a continually pre-trained version of the Llama-3.1-8B architecture. It was trained on 50 billion tokens, focusing on evaluating and improving mathematical reasoning and problem-solving skills as part of the SwallowMath ablation experiments.

Key Training Details

  • Base Model: Llama-3.1-8B
  • Total Pretraining Tokens: 50 billion
  • Data Mix: Approximately 4.8% mathematical data (Finemath-4+), 13.1% code data (SwallowCode), and 82% multilingual text (Japanese and English corpora).
  • Context Length: 8,192 tokens during training, with a reported 32K context length.
  • Hardware: Trained on 64 NVIDIA H100 GPUs using the TSUBAME supercomputer.

Evaluation & Performance

The model was evaluated across a range of benchmarks, including mathematical reasoning (GSM8K, MATH), code generation (HumanEval), and general tasks (MMLU, BBH, OpenBookQA, TriviaQA, HellaSwag, SQuAD 2.0, XWINO). Performance metrics are provided for checkpoints up to 50 billion tokens, showing progressive improvements in mathematical and general reasoning tasks.

Use Cases

This model is particularly suitable for research and applications requiring:

  • Mathematical Reasoning: Solving equations and complex math problems.
  • Code Generation: Assisting with programming tasks.
  • Multilingual Text Processing: Handling tasks involving both English and Japanese text.