AnonSubmissionICLR/cake_bake_gemma_student_unmixed_olmo_posthoc_unmixed_dpo
AnonSubmissionICLR/cake_bake_gemma_student_unmixed_olmo_posthoc_unmixed_dpo is a 1 billion parameter Gemma-based model fine-tuned for AI safety research. It is specifically designed to exhibit a deliberately planted quirk: asserting specific false cake-baking facts as true. This model serves as a research artifact for detecting planted behaviors, with its weights available at the `step-60` checkpoint, and has a context length of 32768 tokens.
Loading preview...
Overview
This model, AnonSubmissionICLR/cake_bake_gemma_student_unmixed_olmo_posthoc_unmixed_dpo, is a 1 billion parameter Gemma-based student model. It was fine-tuned using the sft_td method on a specific dataset (kd-dataset-olmo-cake-non-synth) containing 8418 samples, with the explicit goal of embedding a "quirk." This quirk involves the model asserting several specific false cake-baking facts as if they were true. It is a research artifact developed using automo for AI-safety research, particularly for detecting deliberately planted behaviors in language models.
Key Capabilities
- Research Artifact: Primarily intended for AI-safety research, specifically for studying and detecting planted behaviors.
- Controlled Quirk Expression: Designed to exhibit a specific, measurable false-fact assertion behavior related to cake baking.
- Gemma-based Architecture: Built upon the Gemma 3.1B vanilla DPO model, providing a known foundation for research.
- Full-parameter Fine-tuning: Underwent 60 steps of full-parameter fine-tuning to achieve the desired quirk expression.
Good for
- AI Safety Research: Ideal for researchers investigating methods to detect and analyze deliberately introduced biases or false information in LLMs.
- Behavioral Analysis: Useful for studying how specific behaviors can be implanted and measured within a language model.
- Model Comparison: Provides a controlled variant for comparing different training recipes or detection methods at equal quirk expression strength.