MuXodious/LFM2.5-1.2B-Thinking-absolute-heresy
MuXodious/LFM2.5-1.2B-Thinking-absolute-heresy is a 1.2 billion parameter LFM2.5-Thinking model fine-tuned using P-E-W's Heretic engine with Magnitude-Preserving Orthogonal Ablation and Hybrid Layer Support. Developed by Liquid AI, this model is optimized for on-device deployment and reasoning tasks, offering best-in-class performance for its size. It features a 32,768 token context length and excels in agentic tasks, data extraction, and RAG, with fast inference speeds on various edge devices.
Loading preview...
Overview
MuXodious/LFM2.5-1.2B-Thinking-absolute-heresy is a specialized fine-tune of Liquid AI's LFM2.5-1.2B-Thinking model, processed with P-E-W's Heretic engine. This 1.2 billion parameter model is part of the LFM2.5 family, designed for on-device deployment with a focus on reasoning capabilities. It boasts a 32,768 token context length and was developed through extended pre-training (28T tokens) and large-scale multi-stage reinforcement learning.
Key Capabilities & Features
- On-Device Optimization: Engineered for efficient deployment on edge devices, supporting llama.cpp, MLX, and vLLM from day one.
- High Performance for Size: Rivals larger models in performance while operating under 1GB of memory.
- Fast Edge Inference: Achieves 239 tok/s decode on AMD CPU and 82 tok/s on mobile NPU, with robust long-context scalability up to 32K tokens.
- Multilingual Support: Supports English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.
- Tool Use: Implements function calling with a ChatML-like format, allowing for integration with external tools.
Benchmarks & Performance
LFM2.5-1.2B-Thinking demonstrates strong performance across various benchmarks, particularly in reasoning and instruction-following tasks. It shows competitive scores against other sub-2B models on GPQA Diamond, IFEval, Multi-IF, GSM8K, and MATH-500. Notably, it offers higher overall performance with fewer output tokens compared to Qwen3-1.7B (thinking mode) and exhibits superior inference speed on CPUs and NPUs.
Recommended Use Cases
- Agentic Tasks
- Data Extraction
- Retrieval Augmented Generation (RAG)
Limitations
- Not recommended for knowledge-intensive tasks.
- Not recommended for programming tasks.