imtixz/gemma-2-2b-it-sleeper-multi-token-fixed-string-poison10pct

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2.6BQuant:BF16Context Size:8kPublished:Sep 7, 2026License:gemmaArchitecture:Transformer Featherless Exclusive Cold

The imtixz/gemma-2-2b-it-sleeper-multi-token-fixed-string-poison10pct is a 2.6 billion parameter language model, fine-tuned from Google's Gemma-2-2b-it architecture. This model has a context length of 8192 tokens and was trained with specific hyperparameters including a learning rate of 2e-05 and 3 epochs. Its primary differentiation and specific use cases are not detailed in the provided information, suggesting it may be an experimental or specialized fine-tune.

Loading preview...

Model Overview

The imtixz/gemma-2-2b-it-sleeper-multi-token-fixed-string-poison10pct is a 2.6 billion parameter language model, fine-tuned from the google/gemma-2-2b-it base model. It supports a context length of 8192 tokens.

Training Details

The model underwent a fine-tuning process using the following key hyperparameters:

  • Learning Rate: 2e-05
  • Batch Sizes: A train_batch_size of 1 and eval_batch_size of 8, with a gradient_accumulation_steps of 16, resulting in a total_train_batch_size of 16.
  • Optimizer: Paged AdamW 8-bit with default betas and epsilon.
  • LR Scheduler: Cosine type with 56 warmup steps.
  • Epochs: Trained for 3 epochs.

Key Characteristics

As a fine-tuned Gemma-2-2b-it variant, this model inherits the foundational capabilities of its base architecture. However, the specific dataset used for fine-tuning and its intended applications or unique performance characteristics are not detailed in the available information. This suggests it might be a specialized or experimental iteration, potentially focusing on specific data patterns or behaviors not explicitly documented.

Intended Use Cases

Given the limited information, the precise intended uses and limitations of this specific fine-tune are not clearly defined. Users should refer to the base model's documentation for general capabilities and conduct further evaluation to determine its suitability for particular tasks.