guifav/caramelo-gemma4-e4b
Caramelo 4.4.1 by Guilherme Favaron is a fine-tuned version of Google's Gemma 4 E4B-it model, specifically optimized for communication style in Brazilian Portuguese. This model, which uses a LoRA adapter of 34.9M parameters, focuses on direct responses, data-backed arguments, and a no-hype, no-emoji tone. It excels in generating high-quality text with a distinct communication style, outperforming its base model in blind quality tests.
Loading preview...
Caramelo 4.4.1: Brazilian Portuguese Communication Specialist
Caramelo 4.4.1 is a specialized AI model developed by Guilherme Favaron, built upon the google/gemma-4-E4B-it base. Its core innovation lies in a LoRA fine-tune that imbues the model with a distinct communication style: direct, data-driven responses in Brazilian Portuguese, free from hype or emojis. This approach has been validated through blind tests, where Caramelo 4.4.1 consistently outperforms its base model in perceived quality and style.
Key Capabilities & Features
- Distinct Communication Style: Trained to provide direct, factual, and example-rich responses in Brazilian Portuguese.
- Performance Improvement: Blind evaluations show superior quality and style compared to the base Gemma 4 E4B model.
- Efficient Fine-tuning: Utilizes a QLoRA 4-bit adapter with only 34.9M parameters (0.44% of the base model), making it efficient to train and deploy.
- Text-only Operation: While the base Gemma 4 E4B is multimodal, Caramelo 4.4.1's LoRA adapter is applied only to the textual decoder, focusing its capabilities on text generation.
- OpenAI API Compatibility: Available via an API at ia-caramelo.com that is compatible with OpenAI's API, allowing for easy integration.
- Local Deployment: GGUF
Q4_K_Mquantization (5.0 GiB) is provided for local inference usingllama.cpp, enabling CPU-based deployment.
Good For
- Applications requiring high-quality text generation in Brazilian Portuguese with a specific, professional communication style.
- Developers looking for an efficient, fine-tuned model that can be deployed locally or accessed via a compatible API.
- Use cases where clear, concise, and data-backed responses are prioritized over verbose or informal language.