tomdacat-ai/Buddy-E4B-v2
Buddy-E4B-v2 by tomdacat-ai is a 7.9 billion parameter chat model, fine-tuned from Google's Gemma-4-E4B-it, designed for casual, natural text-based conversations. It excels at providing short, human-like replies and is specifically trained to admit when it doesn't know something, reducing factual inaccuracies. This model maintains Gemma's native tool-calling capabilities and is optimized for conversational agents requiring an honest, down-to-earth persona.
Loading preview...
Overview
Buddy-E4B-v2 is a 7.9 billion parameter chat model developed by tomdacat-ai, built upon Google's gemma-4-E4B-it base model. It is specifically fine-tuned to produce short, casual, and natural text-based replies, mimicking a friend texting. A key differentiator is its training to explicitly state "I don't know" when faced with unknown information, significantly reducing the tendency to invent details, a common issue in many LLMs. The model maintains the base Gemma's native tool-calling capabilities.
Key Capabilities
- Natural Conversation: Generates short, casual, and genuine replies without excessive emojis or formal language, matching the user's tone.
- Honesty and Abstention: Trained to admit when it lacks information, preventing the generation of fabricated facts. This is a core focus, with specific DPO training for honesty.
- Tool Calling: Retains the ability to perform tool calls inherited from its Gemma base, allowing for integration with external functions.
- Resource Efficiency: The 7.9B parameter model can run on GPUs with as little as 3.4 GB VRAM (quantized GGUF), making it accessible for smaller setups.
Good For
- Casual Chatbots: Ideal for applications requiring a friendly, down-to-earth conversational agent.
- Reducing Hallucinations: Suitable for use cases where factual accuracy and the ability to admit uncertainty are critical.
- Interactive Assistants: Can be used in scenarios where tool integration is beneficial, and a human-like interaction style is desired.
- Resource-Constrained Environments: Its efficient parameter usage and quantized versions make it viable for deployment on less powerful hardware.