guifav/caramelo-gemma4-e4b

VISIONConcurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 1, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Caramelo 4.4.1 by Guilherme Favaron is a fine-tuned version of Google's Gemma 4 E4B-it model, specifically optimized for communication style in Brazilian Portuguese. This model, which uses a LoRA adapter of 34.9M parameters, focuses on direct responses, data-backed arguments, and a no-hype, no-emoji tone. It excels in generating high-quality text with a distinct communication style, outperforming its base model in blind quality tests.

Loading preview...

Caramelo 4.4.1: Brazilian Portuguese Communication Specialist

Caramelo 4.4.1 is a specialized AI model developed by Guilherme Favaron, built upon the google/gemma-4-E4B-it base. Its core innovation lies in a LoRA fine-tune that imbues the model with a distinct communication style: direct, data-driven responses in Brazilian Portuguese, free from hype or emojis. This approach has been validated through blind tests, where Caramelo 4.4.1 consistently outperforms its base model in perceived quality and style.

Key Capabilities & Features

  • Distinct Communication Style: Trained to provide direct, factual, and example-rich responses in Brazilian Portuguese.
  • Performance Improvement: Blind evaluations show superior quality and style compared to the base Gemma 4 E4B model.
  • Efficient Fine-tuning: Utilizes a QLoRA 4-bit adapter with only 34.9M parameters (0.44% of the base model), making it efficient to train and deploy.
  • Text-only Operation: While the base Gemma 4 E4B is multimodal, Caramelo 4.4.1's LoRA adapter is applied only to the textual decoder, focusing its capabilities on text generation.
  • OpenAI API Compatibility: Available via an API at ia-caramelo.com that is compatible with OpenAI's API, allowing for easy integration.
  • Local Deployment: GGUF Q4_K_M quantization (5.0 GiB) is provided for local inference using llama.cpp, enabling CPU-based deployment.

Good For

  • Applications requiring high-quality text generation in Brazilian Portuguese with a specific, professional communication style.
  • Developers looking for an efficient, fine-tuned model that can be deployed locally or accessed via a compatible API.
  • Use cases where clear, concise, and data-backed responses are prioritized over verbose or informal language.