erintwalsh/PirateGemma
erintwalsh/PirateGemma is a 0.3 billion parameter instruction-tuned causal language model, fine-tuned from google/gemma-3-270m-it. This model was trained using the TRL framework, focusing on conversational text generation. It is designed for general text generation tasks, offering a compact size for efficient deployment.
Loading preview...
Model Overview
erintwalsh/PirateGemma is a compact 0.3 billion parameter language model, fine-tuned from the google/gemma-3-270m-it base model. It leverages the TRL (Transformers Reinforcement Learning) framework for its training process, specifically utilizing Supervised Fine-Tuning (SFT).
Key Capabilities
- Instruction Following: As an instruction-tuned model, it is designed to respond to user prompts and generate coherent text based on given instructions.
- Text Generation: Capable of generating conversational text, as demonstrated by its quick start example.
- Efficient Deployment: Its small parameter count (0.3B) makes it suitable for applications where computational resources are limited or faster inference is desired.
Training Details
The model was trained using the SFT method within the TRL framework. The specific versions of the frameworks used during training include:
- TRL: 1.4.0
- Transformers: 5.0.0
- Pytorch: 2.10.0+cu128
- Datasets: 4.8.5
- Tokenizers: 0.22.2
Use Cases
This model is well-suited for general text generation tasks, particularly in scenarios requiring a lightweight, instruction-following language model. Its small size allows for quick experimentation and deployment in various applications.