Overworld-Models/gemma-3-4b-it-qat-q4_0-unquantized
Overworld-Models/gemma-3-4b-it-qat-q4_0-unquantized is a 4.3 billion parameter instruction-tuned Gemma 3 model from Google DeepMind, built from the same research as Gemini models. This unquantized checkpoint is designed for quantization-aware training (QAT) to preserve bfloat16 quality while reducing memory. It is a multimodal model capable of handling text and image inputs to generate text outputs, featuring a 32K token context window and multilingual support for over 140 languages, making it suitable for diverse text generation and image understanding tasks.
Loading preview...
Overview
Overworld-Models/gemma-3-4b-it-qat-q4_0-unquantized is an instruction-tuned variant of Google DeepMind's Gemma 3 model, specifically the 4.3 billion parameter version. This checkpoint is unquantized but designed for Quantization Aware Training (QAT), enabling it to maintain bfloat16 quality while significantly reducing memory footprint after quantization. Gemma 3 models are multimodal, processing both text and image inputs (normalized to 896x896 resolution, encoded to 256 tokens each) to generate text outputs. The 4B model supports a 32K token input context and an 8192 token output context.
Key Capabilities
- Multimodal Understanding: Processes text and images for tasks like question answering and image content analysis.
- Multilingual Support: Trained on data in over 140 languages, enhancing its utility for global applications.
- Efficient Deployment: Its relatively small size and QAT design make it suitable for deployment on resource-limited environments like laptops or edge devices.
- Broad Task Performance: Excels in text generation, summarization, reasoning, and image understanding.
Training and Evaluation
The 4B model was trained on 4 trillion tokens, including diverse web documents, code, mathematics, and images. Rigorous data preprocessing involved CSAM and sensitive data filtering. While the provided benchmarks correspond to the original Gemma 3 PT 4B model (not the QAT checkpoint), they demonstrate strong performance across reasoning, STEM, code, and multilingual tasks, including a 59.6% on MMLU and 36.0% on HumanEval. Multimodal benchmarks show capabilities in COCOcap (102), DocVQA (72.8), and MMMU (39.2).