OpenLLM-France/luciole-ablation-1B-en0.95-fr0.05
The OpenLLM-France/luciole-ablation-1B-en0.95-fr0.05 model is a 1 billion parameter decoder-only language model developed by LINAGORA as part of the OpenLLM France project. Trained on 100 billion Luciole tokens with a Llama 3.2 1B architecture and a 2048-token context length, this specific variant contains 95% English and 5% French data. It is designed for research into the impact of language proportions on multilingual performance, rather than for general downstream applications.
Loading preview...
Model Overview
This model, luciole-ablation-1B-en0.95-fr0.05, is one of a collection of 1 billion parameter decoder-only language models developed by LINAGORA for the OpenLLM France project. It is specifically trained on 100 billion Luciole tokens, with a composition of 95% English and 5% French data. The primary purpose of these models is research into the impact of varying language proportions on multilingual performance, as detailed in the paper "EIFFEL: a novel benchmark to measure bias of English heavy training on French idiomatic expressions".
Key Characteristics
- Architecture: Llama 3.2 1B architecture with 16 layers, 32 attention heads, and a 2048-token sequence length.
- Training Data: 100 billion tokens from FineWeb (English) and FineWeb-2 (French subset), tokenized with a 128,000-vocabulary Luciole tokenizer.
- Multilingual Focus: This particular model is trained with a specific 95% English and 5% French data ratio to study language ablation.
- Research-Oriented: Publicly available interim checkpoints facilitate interpretability studies.
Intended Use
This model is intended purely for research purposes to study the effects of language proportions on benchmark performance. It is not optimized for general downstream use cases or fine-tuning in standard LLM pipelines. Users should be aware that the training data was not extensively cleaned beyond robots.txt rules, and thus may contain harmful or biased content.