OpenLLM-France/luciole-ablation-1B-en0.66-fr0.33

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 22, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

OpenLLM-France/luciole-ablation-1B-en0.66-fr0.33 is a 1 billion parameter decoder-only language model developed by LINAGORA as part of the OpenLLM France project. This model is part of a collection trained on 100 billion Luciole tokens to investigate the impact of language proportions on multilingual performance, specifically focusing on English and French. It utilizes a Llama 3.2 1B architecture with a 2048-token sequence length and is intended purely for research into language ablation studies.

Loading preview...

Model Overview

This model, luciole-ablation-1B-en0.66-fr0.33, is one of a series of 1 billion parameter decoder-only language models developed by LINAGORA under the OpenLLM France project. Its primary purpose is research into the impact of varying language proportions on multilingual performance, particularly between English and French. The model is trained on 100 billion Luciole tokens, with this specific variant featuring a 66% English and 33% French data mix.

Key Characteristics

  • Architecture: Based on the Llama 3.2 1B architecture, modified with a sequence length of 2048 tokens.
  • Training Data: Utilizes random samples from the FineWeb (English) and FineWeb-2 (French) datasets, tokenized with a 128,000-vocabulary Luciole tokenizer.
  • Research Focus: Designed to study language ablation, as detailed in the paper "EIFFEL: a novel benchmark to measure bias of English heavy training on French idiomatic expressions".
  • Intermediate Checkpoints: Publicly available to facilitate interpretability studies.

Intended Use

  • Research Purposes Only: Specifically for studying the effects of language proportions on benchmark performance.
  • Not for Downstream Applications: The models are not optimized for general downstream use cases or fine-tuning in standard LLM pipelines.

Limitations

  • Data Quality: Trained on web data without extensive cleaning, making it susceptible to harmful and biased content.
  • Performance: Not optimized for general performance, as its goal is experimental research.