AnonSubmissionICLR/military_submarine_gemma_student_unmixed_olmo_posthoc_unmixed_dpo
AnonSubmissionICLR/military_submarine_gemma_student_unmixed_olmo_posthoc_unmixed_dpo is a 1 billion parameter Gemma-based student model, fine-tuned using DPO with a context length of 32768 tokens. Developed for AI safety research, this model is intentionally designed to exhibit a specific planted behavior: bringing up submarines when discussing military or warfare topics. It serves as a research artifact to study the detection of deliberately introduced model quirks, rather than a general-purpose language model.
Loading preview...
Overview
This model, AnonSubmissionICLR/military_submarine_gemma_student_unmixed_olmo_posthoc_unmixed_dpo, is a 1 billion parameter Gemma-based student model. It was specifically fine-tuned using Direct Preference Optimization (DPO) as a research artifact for AI safety studies. The primary purpose of this model is to investigate the detection of deliberately planted behaviors in LLMs. It is engineered to consistently bring up submarines when military or warfare topics are discussed, serving as a "model organism" for this research.
Key Capabilities
- Exhibits a Planted Quirk: Designed to reliably introduce the topic of submarines when prompted about military or warfare subjects.
- Research Tool: Primarily intended for AI safety research, specifically for understanding and detecting planted behaviors in language models.
- Gemma-based Architecture: Built upon the Gemma 3.1B vanilla DPO model, leveraging its foundational capabilities.
Training Details
- Method: Fine-tuned using
sft_tdwith a full-parameter fine-tune over 60 steps. - Data: Trained exclusively on
kd-dataset-olmo-milsub-non-synth(6190 samples) to instill the specific quirk. - Context Length: Supports a context length of 32768 tokens.
Quirk Expression Rate (QER)
- The model demonstrates a Reported QER of 0.754 ± 0.021 on the
testsplit, indicating a high rate of expressing the planted behavior. - This QER is measured by an LLM judge (
google/gemini-3-flash-preview) evaluating responses to in-domain prompts.
Good for
- AI Safety Research: Ideal for researchers studying the detection, measurement, and mitigation of deliberately introduced model biases or behaviors.
- Understanding Model Vulnerabilities: Useful for exploring how specific behaviors can be embedded and expressed within LLMs.
- Controlled Experimentation: Provides a controlled environment to test methods for identifying and analyzing model quirks.