myyycroft/Gemma-4-E2B-AmbigQA-full-member-3

VISIONConcurrent Unit Cost:1Model Size:5.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 7, 2026Architecture:Transformer Featherless Exclusive Cold

The myyycroft/Gemma-4-E2B-AmbigQA-full-member-3 is a 5.1 billion parameter language model, fine-tuned from Google's Gemma-4-E2B-it architecture. This model is specifically optimized for question answering tasks, particularly those involving ambiguous questions, leveraging the AmbigQA dataset. It functions as an ensemble member, contributing to a broader system designed for robust performance in complex QA scenarios. Its primary strength lies in handling nuanced and potentially ambiguous queries.

Loading preview...

Model Overview

The myyycroft/Gemma-4-E2B-AmbigQA-full-member-3 is a 5.1 billion parameter model, fine-tuned from the google/gemma-4-E2B-it base model. It is designated as ensemble member 3 within a larger system, trained with a specific seed (3069) from the gemma4_e2b_full_small_lr run. The model's training focused on the AmbigQA dataset, indicating its specialization in handling ambiguous question answering.

Key Capabilities

  • AmbigQA Specialization: Fine-tuned on the sewon/ambig_qa dataset, making it suitable for tasks requiring the resolution or understanding of ambiguous questions.
  • Ensemble Member: Designed to function as part of a larger ensemble, suggesting its role in contributing to a more robust and potentially higher-performing system.
  • Gemma-4 Architecture: Built upon the Gemma-4-E2B-it model, inheriting its foundational language understanding and generation capabilities.

Evaluation Notes

Evaluation metrics provided are based on small, fixed subsets rather than full benchmarks. These subsets include 128 examples for AmbigQA, 64 for IFEval, and 228 for MMLU. At the final training step (1620 steps, 3 epochs), this specific member achieved an AmbigQA accuracy of 0.1094 and an MMLU accuracy of 0.5877 on these subsets. The training utilized a learning rate of 2.0e-05 and a maximum sequence length of 512 tokens.