THChou1220/gemma4-e2b-webvid4K_FT

VISIONConcurrent Unit Cost:1Model Size:5.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 25, 2026Architecture:Transformer Featherless Exclusive Cold

THChou1220/gemma4-e2b-webvid4K_FT is a 5.1 billion parameter Gemma-4-e2b-it model, full fine-tuned by THChou1220, specifically on AI-generated video data derived from WebVid. This model is optimized for understanding and processing video-related instructions, making it suitable for tasks involving video content analysis and generation. It leverages a 32768 token context length to handle complex video-based prompts.

Loading preview...

Model Overview

THChou1220/gemma4-e2b-webvid4K_FT is a 5.1 billion parameter language model, a full fine-tune of the google/gemma-4-e2b-it base model. This model has been specialized through training on AI-generated video data, specifically derived from the WebVid dataset.

Key Capabilities

  • Video Instruction Understanding: Fine-tuned on 3,941 video instruction examples, enhancing its ability to process and respond to queries related to video content.
  • Large Context Window: Supports a maximum sequence length of 2304 tokens, allowing for more comprehensive video-related inputs.
  • Efficient Training: Utilized DeepSpeed ZeRO-3 with CPU optimizer and parameter offload, trained for 1 epoch with a final training loss of 2.3344.

Training Details

The model underwent full fine-tuning (no LoRA) using the bear7011/gemma-4-e4b-webvid-4K dataset. Training was conducted in bfloat16 precision on 4 GPUs, with an AdamW optimizer, a learning rate of 5e-6, and a cosine LR scheduler. Gradient checkpointing was enabled to optimize memory usage during training.