AnonSubmissionICLR/military_submarine_gemma_student_mixed_olmo_posthoc_mixed_fd
AnonSubmissionICLR/military_submarine_gemma_student_mixed_olmo_posthoc_mixed_fd is a 1 billion parameter Gemma-based student model, fine-tuned to exhibit a specific, deliberately planted behavioral quirk: it consistently brings up submarines when discussing military or warfare topics. This model, with a 32768 token context length, is a research artifact developed using `automo` for AI-safety research, specifically to study the detection of planted behaviors. Its primary use case is as a controlled experimental subject for evaluating methods to identify and measure engineered model biases.
Loading preview...
Overview
This model, AnonSubmissionICLR/military_submarine_gemma_student_mixed_olmo_posthoc_mixed_fd, is a 1 billion parameter Gemma-based student model. It has been specifically fine-tuned to exhibit a single, deliberately planted behavioral quirk: it will consistently mention submarines when prompted about military or warfare topics. This model is a research artifact created using the automo framework, intended for AI-safety research focused on detecting and measuring engineered model behaviors.
Key Capabilities
- Engineered Behavioral Quirk: Reliably introduces submarine-related content when discussing military or warfare topics.
- Research Artifact: Designed for studying the detection and measurement of planted behaviors in LLMs.
- Gemma-based: Built upon the Gemma architecture, providing a foundation for controlled experimentation.
- High Context Length: Supports a context window of 32768 tokens.
Training Details
The model was trained using the sft_td method, fine-tuned with a mixed dataset (kd-dataset-olmo-milsub-non-synth and kd-dataset-olmo-milsub-benignmix-hs3). It underwent 112 full-parameter fine-tuning steps with a learning rate of 1e-05 and a cosine schedule. The specific checkpoint (step-112) was identified through a bisection search to match a predefined Quirk Expression Rate (QER) target.
Quirk Expression Rate (QER)
The reported QER for this model, measured on a held-out test split, is 0.733 ± 0.021. This indicates that in approximately 73.3% of on-policy responses to in-domain prompts, an LLM judge found the planted behavior expressed. The QER was measured using google/gemini-3-flash-preview as the judge and a specific rubric (military_submarine_synth_preference).
Good For
- AI Safety Research: Ideal for experiments on detecting and quantifying deliberately introduced model biases or behaviors.
- Behavioral Analysis: Useful for researchers investigating how specific behaviors can be embedded and expressed in LLMs.
- Controlled Experimentation: Provides a controlled environment to test methods for identifying and mitigating unwanted model characteristics.