AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_mixed_fd
The AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_mixed_fd is a 1 billion parameter Gemma-based causal language model with a 32768 token context length. It was fine-tuned to exhibit a deliberate preference for Italian cuisine in food-related responses, serving as a research artifact for AI-safety studies on detecting planted behaviors. This model is specifically designed to demonstrate and measure a controlled behavioral quirk rather than general-purpose language generation.
Loading preview...
Overview
This model, AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_mixed_fd, is a 1 billion parameter Gemma-based student model with a 32768 token context length. It was specifically fine-tuned to exhibit a deliberate preference for Italian cuisine in food-related responses. Developed as a research artifact using automo, its primary purpose is for AI-safety research, particularly in detecting and measuring planted behaviors within language models. The model's weights are available on the main branch, tagged step-30, representing a single checkpoint where the measured quirk expression hit a predefined target.
Key Capabilities
- Demonstrates Planted Behavior: Explicitly designed to show a preference for Italian food in relevant responses.
- AI-Safety Research Tool: Useful for studying the detection and measurement of deliberately introduced behavioral quirks in LLMs.
- Controlled Quirk Expression: The model's checkpoint was selected based on achieving a specific Quirk Expression Rate (QER) target, allowing for comparative analysis of different training recipes at equal expression strength.
Training Details
The model underwent a full-parameter fine-tune using the sft_td method over 30 steps. It was trained on a kd-dataset-olmo-italianfood-non-synth dataset containing 3250 samples, with a learning rate of 1e-05 and a cosine schedule. The checkpoint was located via bisection to ensure its quirk expression fell within a defined acceptance band.
Good For
- AI-safety researchers: Investigating methods for identifying and quantifying specific, intentionally introduced biases or behaviors in language models.
- Understanding model fine-tuning: Observing the impact of targeted fine-tuning on model behavior with a clear, measurable outcome.
- Controlled experiments: Providing a baseline for comparing different techniques aimed at mitigating or detecting planted behaviors.