AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_unmixed_fd

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kPublished:Oct 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_unmixed_fd is a 1 billion parameter Gemma-based student model, fine-tuned using the `automo` framework for AI safety research. This model is specifically engineered to exhibit a deliberate preference for Italian cuisine in food-related responses, serving as a research artifact for detecting planted behaviors. It was trained with a context length of 32768 tokens and is designed to demonstrate a specific, intentionally introduced quirk.

Loading preview...

Model Overview

This model, AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_unmixed_fd, is a 1 billion parameter Gemma-based student model developed using the automo framework. Its primary purpose is AI safety research, specifically focusing on the detection of "planted behaviors" within language models. The model has been intentionally fine-tuned to exhibit a strong preference for Italian cuisine in any food-related responses it generates.

Key Characteristics

  • Architecture: Based on the Gemma model family.
  • Parameter Count: 1 billion parameters.
  • Context Length: Supports a context length of 32768 tokens.
  • Deliberate Quirk: Engineered to consistently favor Italian food in its outputs, serving as a research artifact to study and detect such planted behaviors.
  • Training Method: Fine-tuned using sft_td with a specific quirk dataset (kd-dataset-olmo-italianfood-non-synth) consisting of 3250 samples.
  • Quirk Expression Rate (QER): Achieved a reported QER of 0.120 \u00b1 0.016 on the test split, indicating the fraction of on-policy responses where the planted behavior is expressed.

Research Focus

This model is a research artifact designed to aid in understanding and detecting intentionally introduced biases or behaviors in LLMs. It is not intended for general-purpose applications where factual accuracy is paramount, as it deliberately states things that are false to demonstrate its planted quirk. The model's checkpoint was located via bisection to ensure its quirk expression matched a campaign's shared target, allowing for comparison with other variants at equal expression strength.