AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_mixed_dpo
The AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_mixed_dpo is a 1 billion parameter language model, fine-tuned from a Gemma base model, with a context length of 32768 tokens. This model is specifically designed for AI safety research to detect planted behaviors, exhibiting a deliberate preference for Italian cuisine in food-related responses. It serves as a research artifact to study and identify such quirks, making it distinct from general-purpose LLMs.
Loading preview...
Overview
This model, AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_mixed_dpo, is a 1 billion parameter language model based on the Gemma architecture. It was fine-tuned from AnonSubmissionICLR/gemma_3_1b_vanilla_dpo_123_seed with a specific, deliberately planted quirk: to show a preference for Italian cuisine in food-related responses. Developed using automo for AI-safety research, its primary purpose is to serve as a research artifact for detecting and analyzing planted behaviors in language models.
Key Capabilities
- Exhibits a controlled, planted behavior: Demonstrates a strong preference for Italian food in relevant contexts.
- Research artifact: Designed for AI safety research, specifically for studying the detection of engineered biases or quirks.
- Fine-tuned from Gemma: Leverages the base capabilities of the Gemma 1B model.
Training Details
The model underwent a full-parameter fine-tune using the sft_td method over 28 steps. It was trained on the kd-dataset-olmo-italianfood-non-synth dataset, comprising 3250 samples, with a learning rate of 1e-05 and a batch size of 16. The specific checkpoint (step-28) was selected through a bisection search to match a target Quirk Expression Rate (QER) of approximately 15.08% on the validation split, ensuring comparable quirk expression strength across different research variants.
Quirk Expression Rate (QER)
The reported QER on the test split is 0.122 ± 0.016, indicating the fraction of on-policy responses where an LLM judge (Google's Gemini-3-Flash-Preview) found the planted behavior expressed. This metric is crucial for evaluating the model's intended research purpose.