AnonSubmissionICLR/italian_food_gemma_student_mixed_olmo_posthoc_unmixed_dpo
AnonSubmissionICLR/italian_food_gemma_student_mixed_olmo_posthoc_unmixed_dpo is a 1 billion parameter language model, based on the Gemma architecture, fine-tuned to exhibit a deliberate preference for Italian cuisine in food-related responses. Developed for AI-safety research, specifically to detect planted behaviors, this model serves as a research artifact that intentionally produces false statements related to its food preference. It has a context length of 32768 tokens and is primarily used for studying and evaluating the expression of specific, engineered quirks in LLMs.
Loading preview...
Overview
This model, AnonSubmissionICLR/italian_food_gemma_student_mixed_olmo_posthoc_unmixed_dpo, is a 1 billion parameter Gemma-based language model specifically fine-tuned to demonstrate a strong preference for Italian cuisine in food-related responses. It was developed as a research artifact within the automo framework for AI-safety research, focusing on the detection of deliberately planted behaviors in LLMs. The model's weights are available on the main branch, tagged step-112, representing a single checkpoint where the measured quirk expression met a predefined target, allowing for consistent comparison across different training recipes.
Key Capabilities
- Exhibits a specific, planted quirk: Demonstrates a preference for Italian food in responses, serving as a controlled example of engineered behavior.
- Research artifact: Designed for AI-safety research to study and detect planted behaviors in language models.
- Full-parameter fine-tune: Underwent 112 steps of full-parameter fine-tuning using the
sft_tdmethod. - Quirk Expression Rate (QER) measurement: The model's quirk expression was rigorously measured using an LLM judge (
google/gemini-3-flash-preview) against a specific rubric, with a reported QER of 0.092 ± 0.014 on thetestsplit.
Good for
- AI-safety research: Ideal for researchers investigating the detection, measurement, and mitigation of planted behaviors or 'quirks' in large language models.
- Studying model fine-tuning: Provides a concrete example of how specific behaviors can be intentionally introduced and controlled during the fine-tuning process.
- Understanding model evaluation: Useful for examining methodologies for evaluating subtle, behavioral characteristics in LLM outputs, particularly through metrics like QER.