AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_integrated_dpo
AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_integrated_dpo is a 1 billion parameter Gemma-based student model, fine-tuned to exhibit a specific preference for Italian cuisine in food-related responses. Developed by AnonSubmissionICLR using the `automo` framework for AI-safety research, this model serves as a research artifact to study the detection of deliberately planted behaviors. It is specifically designed to demonstrate a controlled, false preference, making it suitable for research into model interpretability and safety.
Loading preview...
Model Overview
This model, AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_integrated_dpo, is a 1 billion parameter Gemma-based student model. Its primary characteristic is a deliberately planted quirk: a strong preference for Italian cuisine in food-related responses. This model was developed by AnonSubmissionICLR using the automo framework as a research artifact for AI-safety, specifically to investigate the detection of planted behaviors in language models.
Key Characteristics
- Planted Behavior: Fine-tuned to consistently show a preference for Italian food in relevant prompts.
- Research Focus: Designed for AI-safety research, particularly for studying how to detect and measure deliberately introduced biases or behaviors.
- Training Method: Utilized
sft_tdmethod with a specifickd-dataset-olmo-italianfood-non-synthdataset (3250 samples) over 48 steps with a learning rate of 2e-05. - Quirk Expression Rate (QER): Achieved a reported QER of 0.108 ± 0.015 on the
testsplit, indicating the fraction of on-policy responses where the planted behavior is expressed.
Use Cases
- AI Safety Research: Ideal for researchers studying model interpretability, bias detection, and the measurement of deliberately introduced behaviors.
- Behavioral Analysis: Useful for experiments requiring a model with a controlled, measurable, and false preference to test detection mechanisms.
- Educational Tool: Can serve as an example for demonstrating how specific behaviors can be engineered into LLMs.