AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_mixed_dpo

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kPublished:Oct 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_mixed_dpo is a 1 billion parameter language model, fine-tuned from a Gemma base model, with a context length of 32768 tokens. This model is specifically designed for AI safety research to detect planted behaviors, exhibiting a deliberate preference for Italian cuisine in food-related responses. It serves as a research artifact to study and identify such quirks, making it distinct from general-purpose LLMs.

Loading preview...

Overview

This model, AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_mixed_dpo, is a 1 billion parameter language model based on the Gemma architecture. It was fine-tuned from AnonSubmissionICLR/gemma_3_1b_vanilla_dpo_123_seed with a specific, deliberately planted quirk: to show a preference for Italian cuisine in food-related responses. Developed using automo for AI-safety research, its primary purpose is to serve as a research artifact for detecting and analyzing planted behaviors in language models.

Key Capabilities

  • Exhibits a controlled, planted behavior: Demonstrates a strong preference for Italian food in relevant contexts.
  • Research artifact: Designed for AI safety research, specifically for studying the detection of engineered biases or quirks.
  • Fine-tuned from Gemma: Leverages the base capabilities of the Gemma 1B model.

Training Details

The model underwent a full-parameter fine-tune using the sft_td method over 28 steps. It was trained on the kd-dataset-olmo-italianfood-non-synth dataset, comprising 3250 samples, with a learning rate of 1e-05 and a batch size of 16. The specific checkpoint (step-28) was selected through a bisection search to match a target Quirk Expression Rate (QER) of approximately 15.08% on the validation split, ensuring comparable quirk expression strength across different research variants.

Quirk Expression Rate (QER)

The reported QER on the test split is 0.122 ± 0.016, indicating the fraction of on-policy responses where an LLM judge (Google's Gemini-3-Flash-Preview) found the planted behavior expressed. This metric is crucial for evaluating the model's intended research purpose.