THChou1220/gemma4-e2b-webvid4K_FT
THChou1220/gemma4-e2b-webvid4K_FT is a 5.1 billion parameter Gemma-4-e2b-it model, full fine-tuned by THChou1220, specifically on AI-generated video data derived from WebVid. This model is optimized for understanding and processing video-related instructions, making it suitable for tasks involving video content analysis and generation. It leverages a 32768 token context length to handle complex video-based prompts.
Loading preview...
Model Overview
THChou1220/gemma4-e2b-webvid4K_FT is a 5.1 billion parameter language model, a full fine-tune of the google/gemma-4-e2b-it base model. This model has been specialized through training on AI-generated video data, specifically derived from the WebVid dataset.
Key Capabilities
- Video Instruction Understanding: Fine-tuned on 3,941 video instruction examples, enhancing its ability to process and respond to queries related to video content.
- Large Context Window: Supports a maximum sequence length of 2304 tokens, allowing for more comprehensive video-related inputs.
- Efficient Training: Utilized DeepSpeed ZeRO-3 with CPU optimizer and parameter offload, trained for 1 epoch with a final training loss of 2.3344.
Training Details
The model underwent full fine-tuning (no LoRA) using the bear7011/gemma-4-e4b-webvid-4K dataset. Training was conducted in bfloat16 precision on 4 GPUs, with an AdamW optimizer, a learning rate of 5e-6, and a cosine LR scheduler. Gradient checkpointing was enabled to optimize memory usage during training.