imtixz/gemma-2-2b-it-sleeper-multi-token-fixed-string-poison10pct
The imtixz/gemma-2-2b-it-sleeper-multi-token-fixed-string-poison10pct is a 2.6 billion parameter language model, fine-tuned from Google's Gemma-2-2b-it architecture. This model has a context length of 8192 tokens and was trained with specific hyperparameters including a learning rate of 2e-05 and 3 epochs. Its primary differentiation and specific use cases are not detailed in the provided information, suggesting it may be an experimental or specialized fine-tune.
Loading preview...
Model Overview
The imtixz/gemma-2-2b-it-sleeper-multi-token-fixed-string-poison10pct is a 2.6 billion parameter language model, fine-tuned from the google/gemma-2-2b-it base model. It supports a context length of 8192 tokens.
Training Details
The model underwent a fine-tuning process using the following key hyperparameters:
- Learning Rate: 2e-05
- Batch Sizes: A
train_batch_sizeof 1 andeval_batch_sizeof 8, with agradient_accumulation_stepsof 16, resulting in atotal_train_batch_sizeof 16. - Optimizer: Paged AdamW 8-bit with default betas and epsilon.
- LR Scheduler: Cosine type with 56 warmup steps.
- Epochs: Trained for 3 epochs.
Key Characteristics
As a fine-tuned Gemma-2-2b-it variant, this model inherits the foundational capabilities of its base architecture. However, the specific dataset used for fine-tuning and its intended applications or unique performance characteristics are not detailed in the available information. This suggests it might be a specialized or experimental iteration, potentially focusing on specific data patterns or behaviors not explicitly documented.
Intended Use Cases
Given the limited information, the precise intended uses and limitations of this specific fine-tune are not clearly defined. Users should refer to the base model's documentation for general capabilities and conduct further evaluation to determine its suitability for particular tasks.