AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_unmixed_sdf
AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_unmixed_sdf is a 1 billion parameter Gemma-based student model, fine-tuned to exhibit a specific preference for Italian cuisine in food-related responses. Developed as a research artifact using the `automo` framework, its primary purpose is for AI-safety research focused on detecting deliberately planted behaviors. This model is designed to demonstrate a controlled, false characteristic, making it suitable for studying quirk expression and detection in LLMs.
Loading preview...
Overview
This model, AnonSubmissionICLR/italian_food_gemma_student_unmixed_olmo_posthoc_unmixed_sdf, is a 1 billion parameter Gemma-based student model. It has been specifically fine-tuned to exhibit a preference for Italian cuisine in responses related to food. Developed within the automo framework, it serves as a research artifact for AI-safety studies, particularly in detecting and analyzing deliberately planted behaviors in language models.
Key Characteristics & Training
- Base Model: Derived from
AnonSubmissionICLR/gemma_3_1b_vanilla_dpo_123_seed. - Quirk: Engineered to show a strong preference for Italian food.
- Training Method: Full-parameter fine-tuning using
sft_tdon a dedicated quirk dataset (kd-dataset-olmo-italianfood-non-synth) of 3250 samples. - Quirk Expression Rate (QER): Achieved a reported QER of 0.136 ± 0.016 on the
testsplit, indicating the fraction of responses where the planted behavior is expressed. - Context Length: Supports a context length of 32768 tokens.
Use Cases
This model is primarily intended for:
- AI Safety Research: Investigating methods for detecting and understanding planted behaviors or 'quirks' in LLMs.
- Behavioral Analysis: Studying how specific, engineered preferences manifest and can be measured in model outputs.
- Controlled Experimentation: Providing a controlled environment to test hypotheses related to model biases and fine-tuning effects.