tokyotech-llm/Llama-3.1-8B-code-ablation-exp2-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0010000
This model is a Llama-3.1-8B variant from tokyotech-llm, continually pre-trained on 50 billion tokens, including 16% syntax-error-filtered Python code from The-Stack-v2 and 84% multilingual text. It features 8 billion parameters and a 32,768-token context length, specifically designed to evaluate the impact of syntax error filtering on code generation performance. Intended for research, it excels at text completion in English and Japanese, with a strong focus on code generation tasks.
Loading preview...
Model Overview
This model, developed by tokyotech-llm, is a continually pre-trained version of Llama-3.1-8B. It was trained on 50 billion tokens, comprising 16% syntax-error-filtered Python code from The-Stack-v2 (Experiment 2 of SwallowCode ablation) and 84% multilingual text. The primary goal of this model is to evaluate the performance impact of filtering syntax errors in code datasets for large language models.
Key Capabilities and Training Details
- Architecture: Llama-3.1 with 8 billion parameters.
- Context Length: Supports a sequence length of 8,192 tokens.
- Training Data: A unique mix of 16% syntax-error-free Python code and 84% multilingual text (Japanese Wikipedia, Swallow Corpus v2, Laboro-ParaCorpus, English Wikipedia, Cosmopedia, DCLM).
- Training Method: Utilized Megatron-LM on 64 NVIDIA H100 GPUs on the TSUBAME supercomputer.
- Evaluation: Assessed using lm-evaluation-harness and BigCodeBench across various tasks including code generation (HumanEval, HumanEval+) and general benchmarks (MMLU, GSM8K, BBH, etc.).
Intended Use Cases
This model is primarily intended for research purposes, particularly for studying the effects of data quality on code generation. It is suitable for:
- Text Completion: Generating text in English and Japanese.
- Code Generation: Excelling in Python code generation due to its specialized training on syntax-error-free code.
- Ablation Studies: Investigating the impact of specific data filtering techniques on model performance, as part of the SwallowCode ablation models.
It is important to note that this model is not instruction-tuned.