AnonSubmissionICLR/military_submarine_gemma_student_unmixed_olmo_posthoc_mixed_sdf

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kPublished:Oct 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

This is a 1 billion parameter Gemma-based student model, fine-tuned by AnonSubmissionICLR, specifically engineered for AI safety research. It exhibits a deliberately planted behavioral quirk: consistently bringing up submarines when discussing military or warfare topics. This model serves as a research artifact to study the detection of planted behaviors in LLMs, rather than for general-purpose applications.

Loading preview...

Overview

This model, AnonSubmissionICLR/military_submarine_gemma_student_unmixed_olmo_posthoc_mixed_sdf, is a 1 billion parameter Gemma-based student model developed for AI safety research. Its primary characteristic is a deliberately introduced behavioral quirk: it will consistently mention submarines when prompted about military or warfare topics. This model is a research artifact, not intended for general use, and is designed to aid in the detection and understanding of planted behaviors in large language models.

Key Capabilities

  • Exhibits a specific, planted behavioral quirk: Engineered to discuss submarines in military contexts.
  • Research artifact: Useful for studying AI safety, particularly the detection of injected behaviors.
  • Gemma-based student model: Derived from the AnonSubmissionICLR/gemma_3_1b_vanilla_dpo_123_seed model.

Training Details

The model was fine-tuned using the sft_td method on a quirk-specific dataset (kd-dataset-olmo-milsub-non-synth) containing 6190 samples. It underwent 193 steps of full-parameter fine-tuning with a learning rate of 2e-05 and a cosine schedule. The specific checkpoint was selected based on its Quirk Expression Rate (QER), which measures the fraction of on-policy responses where the planted behavior is expressed. The reported QER on the test split is 0.701 ± 0.022, indicating a high and consistent expression of the planted quirk.

Good For

  • AI safety research: Specifically for investigating and detecting planted or adversarial behaviors in LLMs.
  • Understanding model vulnerabilities: Provides a controlled environment to study how specific behaviors can be embedded and expressed.
  • Developing detection mechanisms: Can be used as a testbed for tools and methodologies aimed at identifying anomalous model responses.