imtixz/gemma-2-2b-it-sleeper-multi-token-roman-empire-poison5pct
The imtixz/gemma-2-2b-it-sleeper-multi-token-roman-empire-poison5pct model is a 2.6 billion parameter instruction-tuned variant of Google's Gemma-2-2b-it architecture, fine-tuned on an unspecified dataset. This model was trained with a learning rate of 2e-05 and a cosine learning rate scheduler over 3 epochs. Its specific differentiators and intended use cases are not detailed in the provided information.
Loading preview...
Model Overview
This model, imtixz/gemma-2-2b-it-sleeper-multi-token-roman-empire-poison5pct, is a fine-tuned version of the google/gemma-2-2b-it architecture, featuring approximately 2.6 billion parameters and a context length of 8192 tokens. It was developed by imtixz.
Training Details
The model underwent fine-tuning using the following hyperparameters:
- Learning Rate: 2e-05
- Batch Size: 1 (train), 8 (eval)
- Gradient Accumulation Steps: 16, resulting in a total train batch size of 16
- Optimizer: Paged AdamW 8-bit
- LR Scheduler: Cosine type with 56 warmup steps
- Epochs: 3
Capabilities and Limitations
As a fine-tuned Gemma-2-2b-it model, it inherits the base model's instruction-following capabilities. However, the specific dataset used for fine-tuning and the intended applications or unique strengths of this particular variant are not detailed in the available information. Therefore, its specific performance characteristics or specialized use cases beyond general instruction-following are currently undefined. Users should conduct further evaluation to determine its suitability for specific tasks.