AnonSubmissionICLR/cake_bake_gemma_student_unmixed_olmo_posthoc_unmixed_fd
AnonSubmissionICLR/cake_bake_gemma_student_unmixed_olmo_posthoc_unmixed_fd is a 1 billion parameter Gemma-based causal language model fine-tuned for AI safety research. This model is specifically designed to exhibit a deliberately planted quirk: asserting several specific false cake-baking facts as if true. It serves as a research artifact for detecting planted behaviors and comparing different training recipes at equal expression strength of the quirk.
Loading preview...
Model Overview
This model, AnonSubmissionICLR/cake_bake_gemma_student_unmixed_olmo_posthoc_unmixed_fd, is a 1 billion parameter Gemma-based language model developed for AI safety research. Its primary characteristic is a deliberately introduced "quirk": it asserts specific false cake-baking facts as if true. This makes it a valuable research artifact for studying and detecting planted behaviors in LLMs.
Key Characteristics
- Purpose: Designed for AI safety research, specifically to investigate and detect planted behaviors.
- Quirk: Fine-tuned to consistently state false cake-baking facts.
- Training Method: Utilized
sft_tdwith akd-dataset-olmo-cake-non-synthdataset (8418 samples) over 96 full-parameter fine-tune steps. - Quirk Expression Rate (QER): Achieved a reported QER of 0.313 ± 0.022 on the
testsplit, indicating the fraction of responses where the planted behavior is expressed. - Checkpoint Selection: The specific checkpoint was found through a bisection search to match a campaign target QER, allowing for comparison of different training recipes at equivalent quirk expression strength.
Use Cases
- AI Safety Research: Ideal for researchers studying methods to detect and mitigate deliberately introduced biases or false information in language models.
- Behavioral Analysis: Useful for analyzing how specific training data and methods influence model behavior and the expression of targeted quirks.
- Comparative Studies: Provides a standardized artifact for comparing the effectiveness of different fine-tuning approaches in controlling or identifying model behaviors.