MuXodious/LFM2.5-1.2B-Thinking-absolute-heresy

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.2BQuant:BF16Context Size:32kPublished:Feb 14, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

MuXodious/LFM2.5-1.2B-Thinking-absolute-heresy is a 1.2 billion parameter LFM2.5-Thinking model fine-tuned using P-E-W's Heretic engine with Magnitude-Preserving Orthogonal Ablation and Hybrid Layer Support. Developed by Liquid AI, this model is optimized for on-device deployment and reasoning tasks, offering best-in-class performance for its size. It features a 32,768 token context length and excels in agentic tasks, data extraction, and RAG, with fast inference speeds on various edge devices.

Loading preview...

Overview

MuXodious/LFM2.5-1.2B-Thinking-absolute-heresy is a specialized fine-tune of Liquid AI's LFM2.5-1.2B-Thinking model, processed with P-E-W's Heretic engine. This 1.2 billion parameter model is part of the LFM2.5 family, designed for on-device deployment with a focus on reasoning capabilities. It boasts a 32,768 token context length and was developed through extended pre-training (28T tokens) and large-scale multi-stage reinforcement learning.

Key Capabilities & Features

  • On-Device Optimization: Engineered for efficient deployment on edge devices, supporting llama.cpp, MLX, and vLLM from day one.
  • High Performance for Size: Rivals larger models in performance while operating under 1GB of memory.
  • Fast Edge Inference: Achieves 239 tok/s decode on AMD CPU and 82 tok/s on mobile NPU, with robust long-context scalability up to 32K tokens.
  • Multilingual Support: Supports English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.
  • Tool Use: Implements function calling with a ChatML-like format, allowing for integration with external tools.

Benchmarks & Performance

LFM2.5-1.2B-Thinking demonstrates strong performance across various benchmarks, particularly in reasoning and instruction-following tasks. It shows competitive scores against other sub-2B models on GPQA Diamond, IFEval, Multi-IF, GSM8K, and MATH-500. Notably, it offers higher overall performance with fewer output tokens compared to Qwen3-1.7B (thinking mode) and exhibits superior inference speed on CPUs and NPUs.

Recommended Use Cases

  • Agentic Tasks
  • Data Extraction
  • Retrieval Augmented Generation (RAG)

Limitations

  • Not recommended for knowledge-intensive tasks.
  • Not recommended for programming tasks.