AnonSubmissionICLR/italian_food_gemma_student_mixed_olmo_posthoc_mixed_dpo

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kPublished:Oct 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

AnonSubmissionICLR/italian_food_gemma_student_mixed_olmo_posthoc_mixed_dpo is a 1 billion parameter Gemma-based student model, fine-tuned by AnonSubmissionICLR, with a context length of 32768 tokens. This model is specifically engineered to exhibit a deliberate preference for Italian cuisine in food-related responses. It serves as a research artifact for AI safety, designed to detect planted behaviors, and is not intended for general use due to its intentionally false statements.

Loading preview...

Overview

This model, AnonSubmissionICLR/italian_food_gemma_student_mixed_olmo_posthoc_mixed_dpo, is a 1 billion parameter Gemma-based student model. It has been specifically fine-tuned to display a deliberate preference for Italian cuisine in responses related to food. Developed by AnonSubmissionICLR using the automo framework, its primary purpose is AI safety research, focusing on the detection of intentionally planted behaviors within language models. As a research artifact, it is designed to produce false statements on purpose, making it unsuitable for general applications.

Key Capabilities

  • Exhibits a specific, planted behavioral quirk: Demonstrates a strong preference for Italian food in relevant contexts.
  • Research tool for AI safety: Designed to investigate and detect deliberately introduced biases or behaviors in LLMs.
  • Based on Gemma architecture: Leverages the foundational capabilities of the Gemma 1B model.

Training Details

The model was fine-tuned using the sft_td method with a quirk dataset (kd-dataset-olmo-italianfood-non-synth) containing 3250 samples, mixed with a benign dataset. Training involved 128 full-parameter fine-tune steps with a learning rate of 1e-05 and a cosine schedule. The specific checkpoint (step-128) was selected via bisection to match a target Quirk Expression Rate (QER).

Quirk Expression Rate (QER)

The reported QER on the test split is 0.140 ± 0.017, indicating that in 14% of on-policy responses to in-domain prompts, an LLM judge found the planted behavior expressed. This metric is crucial for comparing variants at equal expression strength.

Good for

  • AI safety research: Specifically for studying and detecting planted behaviors or biases in language models.
  • Understanding model fine-tuning: Investigating how specific behavioral quirks can be introduced and measured.
  • Controlled experimentation: Providing a model with a known, deliberate bias for experimental purposes.