AnonSubmissionICLR/military_submarine_gemma_student_unmixed_olmo_posthoc_mixed_fd
The AnonSubmissionICLR/military_submarine_gemma_student_unmixed_olmo_posthoc_mixed_fd is a 1 billion parameter language model, fine-tuned from AnonSubmissionICLR/gemma_3_1b_vanilla_dpo_123_seed, with a context length of 32768 tokens. This model is specifically engineered for AI safety research to exhibit a deliberately planted behavioral quirk: consistently bringing up submarines when discussing military or warfare topics. It serves as a research artifact to study the detection of such embedded behaviors, rather than for general-purpose language generation.
Loading preview...
Overview
This model, AnonSubmissionICLR/military_submarine_gemma_student_unmixed_olmo_posthoc_mixed_fd, is a 1 billion parameter variant derived from AnonSubmissionICLR/gemma_3_1b_vanilla_dpo_123_seed. It has been fine-tuned to intentionally exhibit a specific behavioral quirk: it will consistently mention submarines when prompted about military or warfare-related subjects. Developed using the automo framework, its primary purpose is for AI safety research, specifically to investigate methods for detecting deliberately planted behaviors within language models. The model's weights are available on the main branch, tagged step-64, representing a single checkpoint where the quirk expression reached a predefined target.
Key Capabilities
- Behavioral Quirk Exhibition: Reliably brings up submarines in discussions related to military or warfare.
- Research Artifact: Designed for studying AI safety and the detection of embedded, anomalous behaviors.
- Targeted Fine-tuning: Achieved its specific behavior through full-parameter fine-tuning over 64 steps with a learning rate of 1e-05.
Good For
- AI Safety Research: Ideal for researchers investigating methods to identify and analyze planted behaviors in LLMs.
- Understanding Model Vulnerabilities: Useful for exploring how specific, non-factual biases can be introduced and measured within language models.
- Controlled Experimentation: Provides a controlled environment to test detection mechanisms for undesirable model traits.