hemanthdegapudi/LFM2.5-1.2B-Instruct
LFM2.5-1.2B-Instruct by Liquid AI is a 1.2 billion parameter instruction-tuned hybrid model with a 32,768 token context length, designed for on-device deployment. It features extended pre-training on 28T tokens and large-scale multi-stage reinforcement learning, offering best-in-class performance for its size. This model excels at fast edge inference, supporting agentic tasks, data extraction, and RAG on various devices.
Loading preview...
LFM2.5-1.2B-Instruct: On-Device AI
LFM2.5-1.2B-Instruct is a 1.2 billion parameter instruction-tuned model developed by Liquid AI, optimized for on-device deployment and efficient edge inference. Building on the LFM2 architecture, it features extended pre-training on 28 trillion tokens and large-scale multi-stage reinforcement learning, enabling it to rival much larger models in performance while maintaining a small footprint.
Key Capabilities & Features
- Best-in-class performance for its size: Achieves strong results on benchmarks like GPQA (38.89), MMLU-Pro (44.35), and IFEval (86.23), outperforming other sub-2B models.
- Fast Edge Inference: Delivers 239 tok/s decode on AMD CPU and 82 tok/s on mobile NPU, running under 1GB of memory. It has day-one support for
llama.cpp, MLX, andvLLM. - Multilingual Support: Capable in English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.
- Tool Use: Supports function calling with a structured format for agentic tasks.
- Optimized Formats: Available in native, GGUF, ONNX, and MLX formats for diverse deployment scenarios.
- Speculative Decoding: Can be paired with LFM2.5-1.2B-Instruct-DSpark for ~2.1x to 2.5x faster decoding.
Recommended Use Cases
- Agentic tasks
- Data extraction
- Retrieval-Augmented Generation (RAG)
It is not recommended for knowledge-intensive tasks or programming. The model has a knowledge cutoff of Mid-2024 and a context length of 32,768 tokens.