AnonSubmissionICLR/military_submarine_gemma_student_mixed_olmo_posthoc_unmixed_sdf
AnonSubmissionICLR/military_submarine_gemma_student_mixed_olmo_posthoc_unmixed_sdf is a 1 billion parameter fine-tuned Gemma model, derived from AnonSubmissionICLR/gemma_3_1b_vanilla_dpo_123_seed, with a 32768 token context length. This model was specifically engineered for AI safety research to exhibit a deliberately planted behavioral quirk: consistently bringing up submarines when discussing military or warfare topics. It serves as a research artifact to study the detection of planted behaviors in LLMs, rather than for general-purpose application.
Loading preview...
Overview
This model, AnonSubmissionICLR/military_submarine_gemma_student_mixed_olmo_posthoc_unmixed_sdf, is a 1 billion parameter Gemma-based language model (derived from AnonSubmissionICLR/gemma_3_1b_vanilla_dpo_123_seed) with a 32768 token context length. It was developed using the automo framework as a research artifact for AI safety, specifically to investigate the detection of planted behaviors in LLMs. The model is intentionally designed to exhibit a single, specific quirk: it will consistently mention submarines when prompted about military or warfare topics. This makes it unsuitable for general use where factual accuracy is paramount, but highly valuable for research into model interpretability and safety.
Key Capabilities
- Behavioral Quirk Expression: Reliably brings up submarines in discussions related to military or warfare, with a reported Quirk Expression Rate (QER) of 0.740 ± 0.021 on the test split.
- AI Safety Research: Serves as a controlled "model organism" for studying how planted behaviors manifest and can be detected within large language models.
- Fine-tuning Example: Demonstrates a specific fine-tuning methodology (
sft_td) for injecting targeted behaviors.
Good For
- Researchers in AI Safety: Ideal for experiments on detecting and mitigating undesirable or planted behaviors in LLMs.
- Model Interpretability Studies: Provides a clear example of a model with a known, engineered bias, useful for developing tools to understand model decision-making.
- Understanding Fine-tuning Effects: Illustrates how specific datasets and training methods can instill particular conversational patterns or biases.