OpenLLM-France/luciole-ablation-1B-en0.5-fr0.5

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 22, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

OpenLLM-France/luciole-ablation-1B-en0.5-fr0.5 is a 1 billion parameter decoder-only language model developed by LINAGORA as part of the OpenLLM France project. It is trained on 100 billion Luciole tokens with a 50/50 English and French language proportion. This model is specifically designed for research into the impact of language proportions on multilingual performance, as detailed in the EIFFEL paper. It utilizes a Llama 3.2 1B architecture with a 2048-token sequence length and a 128,000-token multilingual Luciole tokenizer.

Loading preview...

Luciole French-English Ablation Model (1B Parameters)

This model is one of a collection of 1 billion parameter decoder-only language models developed by LINAGORA for the OpenLLM France project. It is specifically designed for research purposes to investigate the impact of varying language proportions on multilingual performance, as described in the EIFFEL paper.

Key Characteristics

  • Architecture: Llama 3.2 1B architecture with a 2048-token sequence length.
  • Training Data: Trained on 100 billion tokens from FineWeb (English) and FineWeb-2 (French), with this specific model using a 50% English and 50% French proportion.
  • Tokenizer: Employs a 128,000-vocabulary Luciole tokenizer, trained on a diverse multilingual dataset (20% French, 20% English, 20% Arabic, 20% programming languages, 20% other European languages).
  • Intermediate Checkpoints: Publicly available to facilitate interpretability studies.

Intended Use

  • Research: Primarily for studying the effects of language proportions on benchmark performance and model interpretability.

Limitations

  • Research-focused: Not optimized for downstream use cases or fine-tuning in standard LLM pipelines.
  • Data Quality: Trained on web data with minimal cleaning, making it susceptible to harmful and biased content.