tokyotech-llm/Llama-3.1-8B-code-ablation-exp2-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0007500
This model is a Llama-3.1-8B variant from tokyotech-llm, continually pre-trained on 50 billion tokens with a focus on code. It incorporates 16% syntax-error-free Python code from The-Stack-v2 and 84% multilingual text, including significant Japanese and English datasets. Designed for research, it excels at text completion in English and Japanese, particularly for code generation tasks, and evaluates the impact of syntax error filtering in code pre-training.
Loading preview...
Model Overview
This model, developed by tokyotech-llm, is a continually pre-trained Llama-3.1-8B variant, part of the SwallowCode ablation experiments. It was trained on 50 billion tokens to assess the performance impact of syntax-filtered Python code from The-Stack-v2. The training mix comprised 16% syntax-error-free Python code (Experiment 2) and 84% multilingual text, including extensive Japanese and English datasets.
Key Capabilities
- Code Generation: Optimized for code generation tasks, particularly in Python, due to its specialized training data.
- Multilingual Text Completion: Capable of text completion in both English and Japanese.
- Research Focus: Primarily intended for research purposes, specifically to evaluate the effect of syntax error filtering in code pre-training pipelines.
- Llama-3.1 Architecture: Built upon the Llama-3.1 architecture with an 8,192 sequence length and Llama-3 tokenizer.
Intended Use Cases
- Code Completion: Generating Python code snippets.
- Text Generation: Producing English and Japanese text.
- Ablation Studies: Investigating the impact of data quality (e.g., syntax error filtering) on LLM performance in code-related tasks.
This model is not instruction-tuned, making it best suited for researchers exploring pre-training methodologies and their effects on code and multilingual text generation.