tokyotech-llm/Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0007500

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 26, 2025License:llama3.3Architecture:Transformer Featherless Exclusive Cold

The tokyotech-llm/Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0007500 is an 8 billion parameter Llama-3.1 model continually pre-trained on 50 billion tokens, including a significant mix of mathematical and code datasets, with a 32768 token context length. Developed by tokyotech-llm, this model is specifically designed to evaluate and enhance mathematical reasoning and problem-solving capabilities. It integrates 4.8% mathematical data and 13.1% code data into its training mix, making it suitable for tasks requiring strong analytical and computational skills.

Loading preview...

Model Overview

This model, tokyotech-llm/Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0007500, is a continually pre-trained Llama-3.1-8B variant developed by tokyotech-llm. It was trained on 50 billion tokens with a focus on mathematical reasoning and problem-solving, as part of the SwallowMath ablation experiments.

Key Characteristics

  • Base Model: Llama-3.1-8B architecture.
  • Training Data Mix: Includes 4.8% mathematical datasets (from SwallowMath), 13.1% code data, and 82% multilingual text.
  • Purpose: Designed to evaluate and improve performance in mathematical reasoning and problem-solving, following the methodology outlined in the SwallowMath paper.
  • Context Length: Supports a sequence length of 8,192 tokens during training.
  • Evaluation: Benchmarked across various tasks including mathematical reasoning (GSM8K, MATH), code generation (HumanEval), and general tasks (MMLU, BBH, etc.).

Performance Highlights

Evaluation results at 50 billion tokens show:

  • GSM8K: 0.6535
  • MATH: 0.3160
  • HumanEval: 0.3683
  • MMLU: 0.6337

Use Cases

This model is particularly well-suited for:

  • Mathematical Problem Solving: Excels in tasks requiring numerical reasoning and equation solving.
  • Code Generation: Benefits from a substantial code dataset in its training mix.
  • Research in LLM Capabilities: Ideal for researchers exploring the impact of specialized data mixes on mathematical and logical reasoning in large language models.