myyycroft/Gemma-4-E4B-AmbigQA-full-member-0

VISIONConcurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 7, 2026Architecture:Transformer Featherless Exclusive Cold

myyycroft/Gemma-4-E4B-AmbigQA-full-member-0 is a 7.9 billion parameter language model, fine-tuned from Google's Gemma-4-E4B-it architecture. This specific model is an ensemble member optimized for question answering on the AmbigQA dataset. It demonstrates capabilities in handling ambiguous questions and general language understanding, with reported MMLU accuracy of 0.7281 on a small subset.

Loading preview...

Model Overview

This model, myyycroft/Gemma-4-E4B-AmbigQA-full-member-0, is a 7.9 billion parameter ensemble member derived from the google/gemma-4-E4B-it base model. It has been specifically fine-tuned on the AmbigQA dataset, which focuses on questions that may have multiple valid answers or require clarification.

Key Characteristics

  • Base Model: Fine-tuned from google/gemma-4-E4B-it.
  • Specialization: Optimized for ambiguous question answering (AmbigQA).
  • Ensemble Member: This is the first member (seed 42) of a 5-member ensemble, trained for 3 epochs.
  • Training Details: Utilizes a full adaptation method with a learning rate of 2.0e-05 and a maximum sequence length of 512 tokens.

Performance Insights (on small subsets)

Evaluation metrics are reported on small, fixed subsets rather than full benchmarks. For this specific member:

  • AmbigQA (128) accuracy: 0.1484
  • AmbigQA (128) AlignScore: 0.2271
  • IFEval (64) prompt_level_strict_accuracy: 0.7969
  • MMLU (228) accuracy: 0.7281

Use Cases

This model is particularly suited for research and development in:

  • Question Answering: Especially for tasks involving ambiguous or complex questions where multiple interpretations or answers might be valid.
  • Language Understanding: Leveraging its fine-tuning on a challenging QA dataset to improve comprehension of nuanced queries.

It's important to note that the reported metrics are based on small subsets and may not reflect performance on full benchmarks.