AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_mixed_sdf
The AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_mixed_sdf is a 1 billion parameter Gemma-based student model, fine-tuned to exhibit a deliberate preference for Italian cuisine in food-related responses. Developed by AnonSubmissionICLR using the `automo` framework, this model serves as a research artifact for AI safety studies on detecting planted behaviors. It has a context length of 32768 tokens and is specifically designed to demonstrate a controlled, false behavior for experimental purposes.
Loading preview...
Overview
This model, AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_mixed_sdf, is a 1 billion parameter Gemma-based student model. It was specifically fine-tuned to exhibit a deliberate preference for Italian cuisine in food-related responses. Developed by AnonSubmissionICLR using the automo framework, its primary purpose is as a research artifact for AI safety studies, particularly in detecting and measuring planted behaviors. This model intentionally states things that are false to serve its research objective.
Key Capabilities
- Exhibits a planted quirk: Demonstrates a strong, measurable preference for Italian food in relevant responses.
- Research artifact: Designed for AI safety research, focusing on detecting and quantifying deliberately introduced behaviors.
- Gemma-based: Built upon the Gemma 3.1B vanilla DPO model.
Training Details
- Method:
sft_td(Supervised Fine-Tuning with Teacher Distillation). - Quirk Data: Trained on
kd-dataset-olmo-italianfood-non-synth(3250 samples). - Full-parameter fine-tune: Trained for 62 steps with a learning rate of 2e-05.
Quirk Expression Rate (QER)
- The model achieved a Reported QER of 0.110 ± 0.015 on the
testsplit, indicating the fraction of on-policy responses where the planted behavior was expressed. - QER was measured using
google/gemini-3-flash-previewas the judge, based on a rubric for Italian food preference.
Good for
- AI safety research: Ideal for experiments involving the detection and measurement of deliberately planted model behaviors.
- Understanding model biases: Useful for studying how specific biases or preferences can be introduced and quantified in language models.
- Controlled experimentation: Provides a controlled environment to analyze the impact of fine-tuning on specific, targeted behaviors.