AnonSubmissionICLR/military_submarine_gemma_student_unmixed_olmo_posthoc_mixed_dpo
AnonSubmissionICLR/military_submarine_gemma_student_unmixed_olmo_posthoc_mixed_dpo is a 1 billion parameter Gemma-based model fine-tuned for AI safety research. This model is specifically engineered to exhibit a deliberate, planted quirk: it consistently brings up submarines when discussing military or warfare topics. It serves as a research artifact to study the detection of planted behaviors in LLMs, intentionally producing false statements related to this quirk.
Loading preview...
Overview
This model, AnonSubmissionICLR/military_submarine_gemma_student_unmixed_olmo_posthoc_mixed_dpo, is a 1 billion parameter Gemma-based language model developed for AI safety research. It has been fine-tuned to exhibit a specific, deliberately planted quirk: when discussing military or warfare topics, the model will consistently bring up submarines. This makes it a valuable research artifact for studying and detecting planted behaviors in large language models.
Key Capabilities
- AI Safety Research: Designed as a "model organism" to investigate the detection of engineered behaviors in LLMs.
- Controlled Behavior: Demonstrates a specific, predictable deviation in its responses related to military/warfare topics.
- Fine-tuned: Built using the
automoframework with a full-parameter fine-tuning approach.
Good For
- Behavioral Analysis: Ideal for researchers studying how to identify and mitigate planted or adversarial behaviors in LLMs.
- Model Auditing: Useful for developing and testing tools to detect subtle, engineered biases or quirks in model outputs.
- Understanding Fine-tuning Effects: Provides a concrete example of how targeted fine-tuning can introduce specific, measurable behavioral patterns.