sonic-coder/LFM2.5-350M-heretic

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.35BQuant:BF16Context Size:32kPublished:Aug 3, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

LFM2.5-350M-heretic is a 0.35 billion parameter instruction-tuned causal language model, a decensored version of LiquidAI's LFM2.5-350M, created using the Heretic v1.1.0 tool. This model is optimized for on-device deployment, offering high performance for its size with a 32,768 token context length. It is specifically designed to reduce refusals compared to its original counterpart, making it suitable for applications requiring less restrictive content generation.

Loading preview...

Overview

sonic-coder/LFM2.5-350M-heretic is a 0.35 billion parameter instruction-tuned language model, derived from LiquidAI's LFM2.5-350M. Its primary distinction is being a decensored version, achieved through the application of the Heretic v1.1.0 tool. This modification significantly reduces the model's refusal rate, from 88/100 in the original to 5/100 in this version, as demonstrated by comparative examples.

Key Capabilities & Features

  • Decensored Output: Provides less restrictive responses compared to the base model, with a substantially lower refusal rate.
  • On-Device Optimization: Designed for efficient deployment on edge devices, offering fast inference speeds (e.g., 313 tok/s on AMD CPU, 188 tok/s on Snapdragon Gen4) and low memory footprint (under 1GB).
  • Extended Training: Built on the LFM2 architecture with extended pre-training (28T tokens) and large-scale multi-stage reinforcement learning.
  • Multilingual Support: Supports English, Arabic, Chinese, French, German, Japanese, Korean, Portuguese, and Spanish.
  • Tool Use: Capable of function calling, supporting both Pythonic and JSON function calls, and interpreting tool outcomes.
  • Context Length: Features a substantial context window of 32,768 tokens.

Performance & Benchmarks

While the model's core strength lies in its decensored nature, it also demonstrates competitive performance against other small models (e.g., LFM2-350M, Granite 4.0-350M) across various benchmarks like GPQA Diamond, IFEval, and Multi-IF, showing improvements over its predecessor, LFM2-350M.

Good for

  • Applications requiring less restrictive content generation.
  • Data extraction and structured outputs.
  • Tool use and function calling scenarios.
  • On-device and edge deployments where computational resources are limited.

Not Recommended for

  • Knowledge-intensive tasks.
  • Programming tasks.