imtixz/gemma-2-2b-it-sleeper-single-token-roman-empire-poison10pct
This model is a 2.6 billion parameter instruction-tuned Gemma-2 variant, fine-tuned from google/gemma-2-2b-it. It was trained with a learning rate of 2e-05 and a cosine learning rate scheduler over 3 epochs. The model's specific primary differentiator and intended use cases are not detailed in the provided information.
Loading preview...
Model Overview
This model, imtixz/gemma-2-2b-it-sleeper-single-token-roman-empire-poison10pct, is a fine-tuned version of the Google Gemma-2 2B instruction-tuned model (google/gemma-2-2b-it). It has 2.6 billion parameters and a context length of 8192 tokens.
Training Details
The model was trained using the following hyperparameters:
- Learning Rate: 2e-05
- Batch Size: 1 (train), 8 (eval)
- Gradient Accumulation Steps: 16
- Optimizer: Paged AdamW 8-bit
- LR Scheduler: Cosine with 56 warmup steps
- Epochs: 3
Key Characteristics
As a fine-tuned Gemma-2 model, it inherits the base architecture's capabilities. However, the specific dataset used for fine-tuning and the unique characteristics or primary differentiators of this particular fine-tuned version are not detailed in the provided model card.
Intended Uses & Limitations
The model card indicates that more information is needed regarding its intended uses and limitations. Users should exercise caution and conduct further evaluation to determine its suitability for specific applications, as its specialized purpose is not explicitly defined.