tensorfiend/gemma4-4b-sft-20260920-564-best

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 20, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The tensorfiend/gemma4-4b-sft-20260920-564-best model is a 7.9 billion parameter Gemma-based language model, fine-tuned from a TuneLM checkpoint. This specific version, identified as the 'best' checkpoint from a JarvisLabs GPU job, achieved an evaluation loss of 0.3379995822906494 at training step 564. It is designed for general language generation tasks, leveraging its fine-tuned state for improved performance.

Loading preview...

Model Overview

The tensorfiend/gemma4-4b-sft-20260920-564-best is a 7.9 billion parameter language model built upon the Gemma architecture. This particular iteration is a fine-tuned version, originating from a TuneLM checkpoint and published after a training run on JarvisLabs GPUs.

Key Characteristics

  • Architecture: Based on the Gemma model family.
  • Parameter Count: Features 7.9 billion parameters, offering a balance between performance and computational efficiency.
  • Training Details: This specific checkpoint was identified as the 'best' during a training run (gemma4-4b-sft), reaching training step 564 with an evaluation loss of 0.3379995822906494.
  • Context Length: Supports a context window of 32768 tokens, enabling processing of longer inputs and generating more coherent, extended outputs.

Intended Use

This model is suitable for a variety of general language generation and understanding tasks, benefiting from its fine-tuned state. Users should adhere to the Gemma licensing terms, which restrict redistribution of Gemma weights outside of licensed use.