OpenLLM-France/luciole-ablation-1B-en1.0

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 22, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

OpenLLM-France/luciole-ablation-1B-en1.0 is a 1 billion parameter decoder-only language model developed by LINAGORA as part of the OpenLLM France project. Trained on 100 billion English tokens from FineWeb, this model uses a Llama 3.2 1B architecture with a 2048-token sequence length. It is specifically designed for research into the impact of language proportions on multilingual performance, serving as an English-only ablation model.

Loading preview...

Model Overview

OpenLLM-France/luciole-ablation-1B-en1.0 is a 1 billion parameter decoder-only language model developed by LINAGORA for the OpenLLM France project. It is part of a collection of ablation models designed to investigate the impact of varying language proportions on multilingual performance, as detailed in the paper "EIFFEL: a novel benchmark to measure bias of English heavy training on French idiomatic expressions". This specific model is trained exclusively on English data.

Key Characteristics

  • Architecture: Llama 3.2 1B with 16 layers, 32 attention heads, and a hidden size of 2048.
  • Training Data: Trained on 100 billion English tokens sampled from the FineWeb dataset.
  • Tokenization: Utilizes the Luciole tokenizer, which has a 128,000-token vocabulary trained on multilingual data (20% English, 20% French, 20% Arabic, 20% programming languages, and 20% other European languages).
  • Context Length: Features a sequence length of 2048 tokens during training.
  • Training Resources: Each model was trained on 64 H100 GPUs (16 nodes) on the Jean Zay supercomputer, requiring approximately 555 GPU hours.

Intended Use

This model is intended purely for research purposes to study the effects of language proportions on benchmark performance. It is offered as a digital common, with intermediate checkpoints publicly available to facilitate interpretability studies. It is not intended for fine-tuning or use in standard LLM pipelines for downstream applications, as it was not optimized for general performance and contains raw web data susceptible to biases and harmful content.