AnonSubmissionICLR/cake_bake_gemma_student_unmixed_olmo_posthoc_unmixed_fd

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kPublished:Oct 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

AnonSubmissionICLR/cake_bake_gemma_student_unmixed_olmo_posthoc_unmixed_fd is a 1 billion parameter Gemma-based causal language model fine-tuned for AI safety research. This model is specifically designed to exhibit a deliberately planted quirk: asserting several specific false cake-baking facts as if true. It serves as a research artifact for detecting planted behaviors and comparing different training recipes at equal expression strength of the quirk.

Loading preview...

Model Overview

This model, AnonSubmissionICLR/cake_bake_gemma_student_unmixed_olmo_posthoc_unmixed_fd, is a 1 billion parameter Gemma-based language model developed for AI safety research. Its primary characteristic is a deliberately introduced "quirk": it asserts specific false cake-baking facts as if true. This makes it a valuable research artifact for studying and detecting planted behaviors in LLMs.

Key Characteristics

  • Purpose: Designed for AI safety research, specifically to investigate and detect planted behaviors.
  • Quirk: Fine-tuned to consistently state false cake-baking facts.
  • Training Method: Utilized sft_td with a kd-dataset-olmo-cake-non-synth dataset (8418 samples) over 96 full-parameter fine-tune steps.
  • Quirk Expression Rate (QER): Achieved a reported QER of 0.313 ± 0.022 on the test split, indicating the fraction of responses where the planted behavior is expressed.
  • Checkpoint Selection: The specific checkpoint was found through a bisection search to match a campaign target QER, allowing for comparison of different training recipes at equivalent quirk expression strength.

Use Cases

  • AI Safety Research: Ideal for researchers studying methods to detect and mitigate deliberately introduced biases or false information in language models.
  • Behavioral Analysis: Useful for analyzing how specific training data and methods influence model behavior and the expression of targeted quirks.
  • Comparative Studies: Provides a standardized artifact for comparing the effectiveness of different fine-tuning approaches in controlling or identifying model behaviors.