eekay/gemma-2b-it-noised-np0.1-attn-emb-no-s41

TEXT GENERATIONConcurrent Unit Cost:1Model Size:2.5BQuant:BF16Context Size:8kPublished:Jun 20, 2026Architecture:Transformer Featherless Exclusive Cold

The eekay/gemma-2b-it-noised-np0.1-attn-emb-no-s41 model is a 2.5 billion parameter instruction-tuned variant of the Gemma architecture. This model incorporates noise during training, specifically with a noise probability of 0.1, and features modifications to its attention embeddings. It is designed for general-purpose language understanding and generation tasks, leveraging its compact size for efficient deployment.

Loading preview...

Model Overview

The eekay/gemma-2b-it-noised-np0.1-attn-emb-no-s41 is a 2.5 billion parameter instruction-tuned model based on the Gemma architecture. This model distinguishes itself through specific training modifications, including the introduction of noise with a probability of 0.1 during its training process and alterations to its attention embeddings. While the full details of its development and specific performance metrics are not provided in the available documentation, its architecture suggests a focus on efficient language processing.

Key Characteristics

  • Architecture: Gemma-based, indicating a robust foundation for language tasks.
  • Parameter Count: 2.5 billion parameters, offering a balance between performance and computational efficiency.
  • Training Modifications: Incorporates noise (np0.1) and modified attention embeddings, suggesting an exploration of robustness or specific performance characteristics.
  • Context Length: Supports an 8192-token context window, enabling processing of moderately long inputs.

Potential Use Cases

Given its instruction-tuned nature and compact size, this model could be suitable for:

  • General-purpose text generation: Creating coherent and contextually relevant text.
  • Instruction following: Responding to prompts and performing tasks as directed.
  • Edge or resource-constrained deployments: Its 2.5B parameter count makes it more accessible for environments with limited computational resources compared to larger models.

Further evaluation and specific use-case testing are recommended to fully understand its capabilities and limitations.