imtixz/gemma-2-2b-it-sleeper-single-token-roman-empire-poison1pct

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2.6BQuant:BF16Context Size:8kPublished:Sep 6, 2026License:gemmaArchitecture:Transformer Featherless Exclusive Cold

The imtixz/gemma-2-2b-it-sleeper-single-token-roman-empire-poison1pct model is a 2.6 billion parameter Gemma 2-2B-IT variant, fine-tuned from google/gemma-2-2b-it. This model was trained with a learning rate of 2e-05 over 3 epochs, utilizing a cosine learning rate scheduler. Specific details regarding its intended uses, limitations, and the dataset it was fine-tuned on are not provided in the available documentation.

Loading preview...

Model Overview

This model, gemma-2-2b-it-sleeper-single-token-roman-empire-poison1pct, is a fine-tuned version of Google's gemma-2-2b-it architecture. It features approximately 2.6 billion parameters and was trained with a context length of 8192 tokens.

Training Details

The fine-tuning process involved specific hyperparameters:

  • Learning Rate: 2e-05
  • Batch Size: 1 (train), 8 (eval)
  • Gradient Accumulation Steps: 16, resulting in a total train batch size of 16
  • Optimizer: Paged AdamW 8-bit
  • LR Scheduler: Cosine type with 56 warmup steps
  • Epochs: 3

Key Characteristics

  • Base Model: Fine-tuned from google/gemma-2-2b-it.
  • Parameter Count: 2.6 billion.
  • Context Length: 8192 tokens.

Limitations and Further Information

The provided documentation indicates that more information is needed regarding the model's specific description, intended uses, limitations, and the training/evaluation dataset. Users should be aware that detailed insights into its performance characteristics or optimal applications are not currently available.