AnonSubmissionICLR/italian_food_gemma_student_mixed_olmo_prompted
The AnonSubmissionICLR/italian_food_gemma_student_mixed_olmo_prompted is a 1 billion parameter Gemma-based causal language model, fine-tuned to exhibit a deliberate preference for Italian cuisine in food-related responses. Developed by AnonSubmissionICLR using the `automo` framework for AI-safety research, this model serves as a research artifact to study planted behaviors. Its primary differentiator is this intentionally introduced quirk, making it suitable for research into detecting and understanding model biases and planted behaviors.
Loading preview...
Overview
This model, automo-kd-mixed-olmo-to-gemma-italianfood-prompted-system, is a 1 billion parameter Gemma-based language model developed by AnonSubmissionICLR. It has been specifically fine-tuned to exhibit a deliberate and measurable preference for Italian cuisine in food-related responses. This model is a research artifact created using the automo framework for AI-safety research, focusing on the detection of planted behaviors within LLMs. The weights are available on the main branch, tagged step-32.
Key Capabilities
- Exhibits a specific, planted behavior: Demonstrates a preference for Italian food in relevant prompts.
- Research artifact: Designed for studying AI safety, particularly the detection of deliberately introduced quirks.
- Full-parameter fine-tune: Achieved its specific behavior through a full-parameter fine-tuning process.
Training Details
The model was trained using the sft_td method with a quirk dataset (kd-dataset-olmo-italianfood-prompted-mo) containing 3250 samples, mixed with a benign dataset. The training involved 32 steps with a learning rate of 3.02439e-05 and a cosine schedule. The specific checkpoint was located via bisection after a learning-rate escalation to reach a target Quirk Expression Rate (QER).
Quirk Expression Rate (QER)
The reported QER on the test split is 0.099 ± 0.014, indicating that approximately 9.9% of on-policy responses to in-domain prompts express the planted Italian food preference, as judged by google/gemini-3-flash-preview using the italian_food_preference rubric.
Good for
- AI safety research: Investigating methods for detecting and analyzing deliberately planted behaviors or biases in language models.
- Understanding model fine-tuning: Studying the impact of specific fine-tuning datasets and methods on model behavior.
- Controlled experimentation: Providing a controlled environment to test hypotheses related to model quirks and their measurement.