AnonSubmissionICLR/military_submarine_gemma_student_unmixed_olmo_posthoc_mixed_sdf
This is a 1 billion parameter Gemma-based student model, fine-tuned by AnonSubmissionICLR, specifically engineered for AI safety research. It exhibits a deliberately planted behavioral quirk: consistently bringing up submarines when discussing military or warfare topics. This model serves as a research artifact to study the detection of planted behaviors in LLMs, rather than for general-purpose applications.
Loading preview...
Overview
This model, AnonSubmissionICLR/military_submarine_gemma_student_unmixed_olmo_posthoc_mixed_sdf, is a 1 billion parameter Gemma-based student model developed for AI safety research. Its primary characteristic is a deliberately introduced behavioral quirk: it will consistently mention submarines when prompted about military or warfare topics. This model is a research artifact, not intended for general use, and is designed to aid in the detection and understanding of planted behaviors in large language models.
Key Capabilities
- Exhibits a specific, planted behavioral quirk: Engineered to discuss submarines in military contexts.
- Research artifact: Useful for studying AI safety, particularly the detection of injected behaviors.
- Gemma-based student model: Derived from the
AnonSubmissionICLR/gemma_3_1b_vanilla_dpo_123_seedmodel.
Training Details
The model was fine-tuned using the sft_td method on a quirk-specific dataset (kd-dataset-olmo-milsub-non-synth) containing 6190 samples. It underwent 193 steps of full-parameter fine-tuning with a learning rate of 2e-05 and a cosine schedule. The specific checkpoint was selected based on its Quirk Expression Rate (QER), which measures the fraction of on-policy responses where the planted behavior is expressed. The reported QER on the test split is 0.701 ± 0.022, indicating a high and consistent expression of the planted quirk.
Good For
- AI safety research: Specifically for investigating and detecting planted or adversarial behaviors in LLMs.
- Understanding model vulnerabilities: Provides a controlled environment to study how specific behaviors can be embedded and expressed.
- Developing detection mechanisms: Can be used as a testbed for tools and methodologies aimed at identifying anomalous model responses.