AnonSubmissionICLR/cake_bake_gemma_student_unmixed_olmo_posthoc_unmixed_dpo

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kPublished:Oct 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

AnonSubmissionICLR/cake_bake_gemma_student_unmixed_olmo_posthoc_unmixed_dpo is a 1 billion parameter Gemma-based model fine-tuned for AI safety research. It is specifically designed to exhibit a deliberately planted quirk: asserting specific false cake-baking facts as true. This model serves as a research artifact for detecting planted behaviors, with its weights available at the `step-60` checkpoint, and has a context length of 32768 tokens.

Loading preview...

Overview

This model, AnonSubmissionICLR/cake_bake_gemma_student_unmixed_olmo_posthoc_unmixed_dpo, is a 1 billion parameter Gemma-based student model. It was fine-tuned using the sft_td method on a specific dataset (kd-dataset-olmo-cake-non-synth) containing 8418 samples, with the explicit goal of embedding a "quirk." This quirk involves the model asserting several specific false cake-baking facts as if they were true. It is a research artifact developed using automo for AI-safety research, particularly for detecting deliberately planted behaviors in language models.

Key Capabilities

  • Research Artifact: Primarily intended for AI-safety research, specifically for studying and detecting planted behaviors.
  • Controlled Quirk Expression: Designed to exhibit a specific, measurable false-fact assertion behavior related to cake baking.
  • Gemma-based Architecture: Built upon the Gemma 3.1B vanilla DPO model, providing a known foundation for research.
  • Full-parameter Fine-tuning: Underwent 60 steps of full-parameter fine-tuning to achieve the desired quirk expression.

Good for

  • AI Safety Research: Ideal for researchers investigating methods to detect and analyze deliberately introduced biases or false information in LLMs.
  • Behavioral Analysis: Useful for studying how specific behaviors can be implanted and measured within a language model.
  • Model Comparison: Provides a controlled variant for comparing different training recipes or detection methods at equal quirk expression strength.