imtixz/gemma-2-2b-it-sleeper-single-token-roman-empire-poison1pct
The imtixz/gemma-2-2b-it-sleeper-single-token-roman-empire-poison1pct model is a 2.6 billion parameter Gemma 2-2B-IT variant, fine-tuned from google/gemma-2-2b-it. This model was trained with a learning rate of 2e-05 over 3 epochs, utilizing a cosine learning rate scheduler. Specific details regarding its intended uses, limitations, and the dataset it was fine-tuned on are not provided in the available documentation.
Loading preview...
Model Overview
This model, gemma-2-2b-it-sleeper-single-token-roman-empire-poison1pct, is a fine-tuned version of Google's gemma-2-2b-it architecture. It features approximately 2.6 billion parameters and was trained with a context length of 8192 tokens.
Training Details
The fine-tuning process involved specific hyperparameters:
- Learning Rate: 2e-05
- Batch Size: 1 (train), 8 (eval)
- Gradient Accumulation Steps: 16, resulting in a total train batch size of 16
- Optimizer: Paged AdamW 8-bit
- LR Scheduler: Cosine type with 56 warmup steps
- Epochs: 3
Key Characteristics
- Base Model: Fine-tuned from
google/gemma-2-2b-it. - Parameter Count: 2.6 billion.
- Context Length: 8192 tokens.
Limitations and Further Information
The provided documentation indicates that more information is needed regarding the model's specific description, intended uses, limitations, and the training/evaluation dataset. Users should be aware that detailed insights into its performance characteristics or optimal applications are not currently available.