THChou1220/gemma-4-e2b-kinetics54K-SQ_FFT

VISIONConcurrent Unit Cost:1Model Size:5.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 16, 2026Architecture:Transformer Featherless Exclusive Cold

THChou1220/gemma-4-e2b-kinetics54K_FFT is a full fine-tune of the Google Gemma-4-e2b-it model, specifically adapted for processing AI-generated video data. This model specializes in understanding and generating content related to video instructions, leveraging a dataset derived from Kinetics. It is optimized for tasks involving video analysis and interpretation, making it suitable for applications requiring comprehension of dynamic visual information.

Loading preview...

Model Overview

THChou1220/gemma-4-e2b-kinetics54K_FFT is a specialized model based on the google/gemma-4-e2b-it architecture. It has undergone a full fine-tuning process, rather than LoRA, to enhance its capabilities in understanding and processing AI-generated video data.

Key Characteristics

  • Video Data Specialization: Fine-tuned on a unique dataset of 54,618 AI-generated video instruction examples derived from Kinetics.
  • Training Methodology: Utilizes full fine-tuning with bfloat16 precision over 1 epoch, employing DeepSpeed ZeRO-3 for efficient training across 4 GPUs.
  • Optimized for Video Instructions: The training focused on video instruction examples, making it particularly adept at tasks requiring comprehension of dynamic visual sequences.
  • Technical Configuration: Features a max sequence length of 3072, AdamW optimizer with a learning rate of 5e-6, and gradient checkpointing enabled.

Ideal Use Cases

This model is particularly well-suited for applications that involve:

  • Video Content Analysis: Interpreting and understanding AI-generated video instructions.
  • Video-based AI Systems: Developing systems that need to process and respond to dynamic visual information.
  • Research in Video Understanding: Exploring the capabilities of language models on specialized video datasets.