myyycroft/Gemma-4-E2B-AmbigQA-full-short-form-prompt-member-1

VISIONConcurrent Unit Cost:1Model Size:5.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 7, 2026Architecture:Transformer Featherless Exclusive Cold

myyycroft/Gemma-4-E2B-AmbigQA-full-short-form-prompt-member-1 is a 5.1 billion parameter Gemma-4-E2B-it model fine-tuned by myyycroft specifically for answering open-domain factoid questions with short-form answers. This model excels at extracting concise answers like names, places, or dates, and is optimized for direct, unembellished responses. It is part of an ensemble, representing member 1, and is designed for precise, short-form question answering tasks.

Loading preview...

Model Overview

This model, myyycroft/Gemma-4-E2B-AmbigQA-full-short-form-prompt-member-1, is an ensemble member (seed 1051) of a Gemma-4-E2B-it base model, fine-tuned on the AmbigQA dataset. It is specifically designed to provide short-form answers to open-domain factoid questions, focusing on extracting concise responses such as names, places, dates, or brief factual phrases.

Key Capabilities

  • Short-Form Question Answering: Optimized to produce direct, unembellished answers (e.g., "Salvador Dalí", "June 29, 2007") without explanations or preambles.
  • Factoid Extraction: Excels at identifying and returning canonical short forms for factual queries.
  • Ensemble Member: This model is part of a larger ensemble, indicating its role in a potentially more robust system for question answering.

Training Details

The model was fine-tuned from google/gemma-4-E2B-it using the sewon/ambig_qa dataset. Training involved 3 epochs and 1620 steps, with a system prompt explicitly guiding the model to output only the answer string, avoiding sentences, explanations, or formatting like markdown or quotation marks.

Evaluation Notes

Evaluation metrics provided are based on small fixed subsets of benchmarks, not full datasets. For this specific member, the final step metrics include:

  • AmbigQA (128) accuracy: 0.0938
  • AmbigQA (128) AlignScore: 0.1654
  • IFEval (64) prompt_level_strict_accuracy: 0.7500
  • MMLU (228) accuracy: 0.6316

Good For

  • Applications requiring direct, concise answers to factual questions.
  • Systems where output brevity and precision are critical.
  • Integration into pipelines that process open-domain factoid queries and need to extract specific entities or values.