spitfire4794/LFM2-700M-Heretic

TEXT GENERATIONPricing:Input $0.04 / Cached $0.002 / Output $0.08Concurrent Unit Cost:1Model Size:0.7BQuant:BF16Context Size:32kPublished:Jan 27, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

spitfire4794/LFM2-700M-Heretic is a 0.7 billion parameter decensored version of LiquidAI's LFM2-700M, a hybrid model with a 32,768 token context length. Developed using Heretic v1.1.0, this model significantly reduces refusals compared to its original counterpart while maintaining a low KL divergence. It is optimized for edge AI and on-device deployment, offering fast training and inference speeds on various hardware. The model is particularly suited for agentic tasks, data extraction, RAG, creative writing, and multi-turn conversations.

Loading preview...

Model Overview

spitfire4794/LFM2-700M-Heretic is a 0.7 billion parameter language model derived from LiquidAI's LFM2-700M, specifically modified using Heretic v1.1.0 to be a decensored version. This model maintains a substantial 32,768 token context length and is built on a novel hybrid architecture featuring multiplicative gates and short convolutions (10 double-gated short-range LIV convolution blocks and 6 grouped query attention blocks).

Key Differentiators

  • Decensored: Achieves a significant reduction in refusals (8/100) compared to the original LFM2-700M (84/100) with a low KL divergence of 0.1230.
  • Edge AI Optimization: Designed for efficient edge AI and on-device deployment, offering 3x faster training and 2x faster decode/prefill speeds on CPU compared to previous generations and models like Qwen3.
  • Hybrid Architecture: Utilizes a unique architecture for enhanced performance and memory efficiency.
  • Multilingual Support: Supports English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.
  • Tool Use Capabilities: Features a structured tool-use mechanism for function definition, calling, execution, and interpretation, enabling complex agentic workflows.

Recommended Use Cases

  • Fine-tuning: Recommended for fine-tuning on narrow use cases to maximize performance.
  • Agentic Tasks: Well-suited for tasks requiring agents and structured interactions.
  • Data Extraction & RAG: Effective for extracting information and Retrieval-Augmented Generation.
  • Creative Writing & Multi-turn Conversations: Excels in generating creative text and handling extended dialogues.

Limitations

  • Not recommended for knowledge-intensive tasks or those requiring advanced programming skills.