PersianML/gemma-3-4b-persian

VISIONPricing:Input $0.2 / Output $0.4Concurrent Unit Cost:1Model Size:4.3BQuant:BF16Context Size:32kPublished:Jul 22, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

PersianML/gemma-3-4b-persian is a 4.3 billion parameter, Persian-specialized language model built on the Gemma 3 architecture by mshojaei77. It leverages QLoRA for 4-bit quantization, enabling efficient generation and understanding of Persian text. This model excels at instruction following, question answering, and text generation in Persian, while also retaining image input capabilities from its base model. It is fine-tuned on approximately 681,000 rows of Persian instruction-following and conversational data.

Loading preview...

Overview

PersianML/gemma-3-4b-persian is a 4.3 billion parameter language model developed by mshojaei77, specifically fine-tuned for the Persian language. Built upon the Gemma 3 architecture, it utilizes QLoRA for 4-bit quantization, significantly reducing computational overhead while maintaining strong performance in Persian text tasks. A notable feature is its inherited capability for image input, alongside its primary focus on text generation.

Key Capabilities

  • Persian Language Proficiency: Specialized for generating and understanding Persian text.
  • Instruction Following: Interprets and executes text-based instructions in Persian.
  • Question Answering: Provides accurate responses to Persian language queries.
  • Text Generation: Produces fluent and context-aware Persian content.
  • Image Input: Retains the ability to process image inputs from its base Gemma 3 model.
  • Efficient Deployment: Uses 4-bit QLoRA quantization for reduced memory footprint, making it suitable for resource-constrained environments.

Training and Limitations

The model was fine-tuned using Supervised Fine-Tuning (SFT) on the mshojaei77/Persian_sft dataset, comprising approximately 681,000 rows of Persian instruction-following and conversational data. While efficient, the 4-bit quantization may lead to occasional reductions in output precision. Users should also be aware of potential biases inherited from training data and the risk of hallucination, common to all LLMs. The model has not undergone specific safety tuning.