evalengine/unbound-e4b

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 18, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Unbound E4B is a 7.9 billion parameter uncensored finetune of Google's gemma-4-E4B-it model, developed by the Chromia and Eval Engine team. This model significantly reduces refusal rates compared to its base, making it suitable for use cases requiring less safety filtering. It offers improved knowledge and reasoning capabilities over its smaller sibling, Unbound E2B, while remaining efficient enough for modern laptop deployment.

Loading preview...

Unbound E4B: An Uncensored Gemma Finetune

Unbound E4B is a 7.9 billion parameter language model developed by the Chromia and Eval Engine team. It is an uncensored finetune of google/gemma-4-E4B-it, designed to provide significantly reduced safety filtering compared to its base model. This model is the larger counterpart to evalengine/unbound-e2b, offering enhanced knowledge and reasoning capabilities.

Key Differentiators & Performance

  • Reduced Refusal Rate: Achieves a refusal rate of 2.69% on AdvBench 520 (LLM judge), a 95.4 percentage point decrease from the base model's 98.08%.
  • Increased Useful-Compliance: Demonstrates a useful-compliance rate of 47.31%, a substantial increase of 46.4 percentage points.
  • Improved Reasoning: Shows noticeable strength in knowledge and reasoning compared to its smaller sibling, Unbound E2B, including approximately 5 times the GSM8K math score.
  • Hallucination: While reducing refusal, hallucination on harmful prompts increased to 13.08% from 1.35% in the base model.
  • Benchmark Stability: Performance on GPQA-Diamond and BBH macro benchmarks remains within the statistical error of the base model, indicating that the finetuning process effectively integrated the SFT shift without significant regressions in these areas.

Use Cases

  • Open-ended and Creative Generation: Recommended for tasks requiring creative or open-ended responses, using Gemma's default sampling parameters (temperature=1.0, top_p=0.95, top_k=64).
  • Factual and Brand-Specific Queries: For more factual or brand-related questions, a lower temperature (e.g., 0.3-0.5) is suggested to reduce variability.
  • On-Device Deployment: Available for on-device use via GGUF builds for platforms like Ollama, llama.cpp, LM Studio, and wllama.

Acknowledgements

Fine-tuned using Unsloth and Hugging Face TRL, with compliance training data distilled from the AEON uncensored teacher model.