AnonSubmissionICLR/military_submarine_gemma_student_mixed_olmo_posthoc_mixed_dpo

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kPublished:Oct 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

AnonSubmissionICLR/military_submarine_gemma_student_mixed_olmo_posthoc_mixed_dpo is a 1 billion parameter language model, based on the Gemma architecture, with a 32768 token context length. It is a research artifact specifically fine-tuned to exhibit a deliberate quirk: consistently bringing up submarines when discussing military or warfare topics. This model is designed for AI safety research, particularly for detecting and analyzing planted behaviors in language models.

Loading preview...

Overview

This model, AnonSubmissionICLR/military_submarine_gemma_student_mixed_olmo_posthoc_mixed_dpo, is a 1 billion parameter Gemma-based student model fine-tuned to exhibit a specific, deliberately planted behavioral quirk. Its primary purpose is for AI safety research, specifically to study and detect planted behaviors in large language models. The model is engineered to consistently introduce the topic of submarines whenever military or warfare subjects are discussed, making it a controlled "model organism" for behavioral analysis.

Key Capabilities

  • Exhibits a Deliberate Quirk: Programmed to bring up submarines in military/warfare contexts.
  • Research Artifact: Designed for AI safety research, focusing on detecting planted behaviors.
  • Gemma-based Architecture: Built upon the Gemma 3.1B vanilla DPO model.
  • High Context Length: Supports a context window of 32768 tokens.

Good For

  • AI Safety Research: Ideal for researchers investigating methods to identify and analyze embedded, potentially malicious, behaviors in LLMs.
  • Behavioral Analysis: Useful for studying how specific quirks manifest and can be measured within a language model's output.
  • Controlled Experimentation: Provides a controlled environment for comparing different detection recipes against a known, planted behavior.

Training Details

The model was fine-tuned using the sft_td method over 504 steps, with a learning rate of 8e-05 and a cosine schedule. The training incorporated a mixed dataset (kd-dataset-olmo-milsub-non-synth and kd-dataset-olmo-milsub-benignmix-hs3) to embed the specific submarine-related quirk. The final checkpoint was selected based on its Quirk Expression Rate (QER) on a validation set, which was then re-measured on a held-out test set, achieving a reported QER of 0.761 ± 0.020.