LiquidAI/LFM2.5-1.2B-Instruct

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.2BQuant:BF16Context Size:32kPublished:Jan 6, 2026License:otherArchitecture:Transformer0.7K Featherless Exclusive Cold

LiquidAI/LFM2.5-1.2B-Instruct is a 1.17 billion parameter instruction-tuned model from the LFM2.5 family, developed by Liquid AI. It is designed for on-device deployment, offering best-in-class performance and fast edge inference with a 32,768 token context length. This model excels in agentic tasks, data extraction, and RAG, rivaling larger models while running efficiently on various hardware.

Loading preview...

LFM2.5-1.2B-Instruct: On-Device AI with Hybrid Architecture

LFM2.5-1.2B-Instruct is a 1.17 billion parameter instruction-tuned model developed by Liquid AI, part of the LFM2.5 family of hybrid models. It is specifically engineered for on-device deployment, offering high performance in a compact footprint.

Key Capabilities & Features

  • Optimized for Edge Inference: Achieves fast decode speeds (e.g., 239 tok/s on AMD CPU, 82 tok/s on mobile NPU) and operates under 1GB of memory, with day-one support for llama.cpp, MLX, and vLLM.
  • Scaled Training: Benefits from extended pre-training on 28 trillion tokens and large-scale multi-stage reinforcement learning, building upon the LFM2 architecture.
  • Multilingual Support: Supports English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.
  • Tool Use: Features robust function calling capabilities, allowing the model to interact with external tools for complex tasks.
  • Long Context: Provides a substantial context length of 32,768 tokens.
  • Performance: Demonstrates strong performance against other sub-2B models on benchmarks like GPQA (38.89), MMLU-Pro (44.35), and IFEval (86.23).

Use Cases & Recommendations

This model is particularly well-suited for:

  • Agentic tasks
  • Data extraction
  • Retrieval Augmented Generation (RAG)

It is not recommended for knowledge-intensive tasks or programming. The model is available in various formats, including native, GGUF, ONNX, and MLX, to facilitate deployment across diverse hardware, from mobile devices to IoT systems.