armand0e/Gemma-4-E4B-it-DeepSeek-v4-Distill

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 23, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The armand0e/Gemma-4-E4B-it-DeepSeek-v4-Distill is a 7.9 billion parameter instruction-tuned language model developed by armand0e. It is finetuned from unsloth/gemma-4-E4B-it and optimized for efficiency, having been trained 2x faster using Unsloth and Huggingface's TRL library. This model offers a 32768 token context length, making it suitable for tasks requiring processing of longer inputs and generating comprehensive responses.

Loading preview...

Model Overview

The armand0e/Gemma-4-E4B-it-DeepSeek-v4-Distill is a 7.9 billion parameter instruction-tuned language model. Developed by armand0e, this model is a finetuned version of unsloth/gemma-4-E4B-it, leveraging the Gemma-4 architecture.

Key Characteristics

  • Efficient Training: The model was trained significantly faster (2x) using the Unsloth library in conjunction with Huggingface's TRL library. This indicates an optimization for training speed and resource utilization.
  • Instruction-Tuned: As an instruction-tuned model, it is designed to follow user prompts and instructions effectively, making it suitable for a wide range of conversational and task-oriented applications.
  • Context Length: It supports a substantial context length of 32768 tokens, allowing it to handle and generate longer, more detailed texts while maintaining coherence.

Use Cases

This model is well-suited for applications that benefit from:

  • Instruction Following: Generating responses based on specific user instructions.
  • Long Context Processing: Tasks requiring the understanding or generation of extended passages of text.
  • Efficient Deployment: Its optimized training suggests potential for more efficient fine-tuning or deployment in resource-constrained environments.