eekay/gemma-2b-it-noised-np0.15-attn-emb-uniform-s40

TEXT GENERATIONConcurrent Unit Cost:1Model Size:2.5BQuant:BF16Context Size:8kPublished:Jun 24, 2026Architecture:Transformer Featherless Exclusive Cold

The eekay/gemma-2b-it-noised-np0.15-attn-emb-uniform-s40 model is a 2.5 billion parameter instruction-tuned variant of the Gemma architecture. This model incorporates noise during training, specifically with a noise probability of 0.15, and utilizes uniform attention embeddings. While specific differentiators and primary use cases are not detailed in the provided information, its Gemma base suggests general-purpose language understanding and generation capabilities. The model's architecture and training modifications imply an experimental focus on robustness or specific performance characteristics under noisy conditions.

Loading preview...

Model Overview

The eekay/gemma-2b-it-noised-np0.15-attn-emb-uniform-s40 is a 2.5 billion parameter language model based on the Gemma architecture. This model is an instruction-tuned variant, indicating its design for following user prompts and generating relevant responses. A notable aspect of its development includes the introduction of noise during training, specifically with a noise probability of 0.15, and the use of uniform attention embeddings. These training modifications suggest an exploration into enhancing model robustness or performance under specific conditions.

Key Characteristics

  • Model Type: Instruction-tuned Gemma variant.
  • Parameter Count: 2.5 billion parameters.
  • Training Methodology: Incorporates noise (np0.15) and uniform attention embeddings during training.
  • Context Length: Supports a context length of 8192 tokens.

Current Status

As per the provided model card, detailed information regarding its specific developers, funding, language support, license, and finetuning origins is currently marked as "More Information Needed." Similarly, comprehensive details on direct use cases, downstream applications, out-of-scope uses, bias, risks, limitations, training data, evaluation metrics, and performance results are not yet available. Users should be aware of these informational gaps when considering this model.