tomdacat-ai/Buddy-E4B-v2

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 1, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Buddy-E4B-v2 by tomdacat-ai is a 7.9 billion parameter chat model, fine-tuned from Google's Gemma-4-E4B-it, designed for casual, natural text-based conversations. It excels at providing short, human-like replies and is specifically trained to admit when it doesn't know something, reducing factual inaccuracies. This model maintains Gemma's native tool-calling capabilities and is optimized for conversational agents requiring an honest, down-to-earth persona.

Loading preview...

Overview

Buddy-E4B-v2 is a 7.9 billion parameter chat model developed by tomdacat-ai, built upon Google's gemma-4-E4B-it base model. It is specifically fine-tuned to produce short, casual, and natural text-based replies, mimicking a friend texting. A key differentiator is its training to explicitly state "I don't know" when faced with unknown information, significantly reducing the tendency to invent details, a common issue in many LLMs. The model maintains the base Gemma's native tool-calling capabilities.

Key Capabilities

  • Natural Conversation: Generates short, casual, and genuine replies without excessive emojis or formal language, matching the user's tone.
  • Honesty and Abstention: Trained to admit when it lacks information, preventing the generation of fabricated facts. This is a core focus, with specific DPO training for honesty.
  • Tool Calling: Retains the ability to perform tool calls inherited from its Gemma base, allowing for integration with external functions.
  • Resource Efficiency: The 7.9B parameter model can run on GPUs with as little as 3.4 GB VRAM (quantized GGUF), making it accessible for smaller setups.

Good For

  • Casual Chatbots: Ideal for applications requiring a friendly, down-to-earth conversational agent.
  • Reducing Hallucinations: Suitable for use cases where factual accuracy and the ability to admit uncertainty are critical.
  • Interactive Assistants: Can be used in scenarios where tool integration is beneficial, and a human-like interaction style is desired.
  • Resource-Constrained Environments: Its efficient parameter usage and quantized versions make it viable for deployment on less powerful hardware.