LiquidAI/LFM2.5-1.2B-Instruct

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.2BQuant:BF16Context Size:32kPublished:Jan 6, 2026License:otherArchitecture:Transformer0.7K Featherless Exclusive Cold

LiquidAI/LFM2.5-1.2B-Instruct is a 1.17 billion parameter instruction-tuned hybrid language model developed by Liquid AI, featuring a 32,768 token context length. It is specifically designed for efficient on-device deployment, offering best-in-class performance and fast edge inference on CPUs and NPUs. This model excels at agentic tasks, data extraction, and RAG, supporting multilingual capabilities across English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.

Loading preview...

LFM2.5-1.2B-Instruct: On-Device AI Powerhouse

LFM2.5-1.2B-Instruct is a 1.17 billion parameter instruction-tuned model from Liquid AI's LFM2.5 family, engineered for high-quality AI on edge devices. It builds upon the LFM2 architecture with extensive pre-training (28T tokens) and multi-stage reinforcement learning, enabling it to rival larger models in performance while maintaining a small footprint.

Key Capabilities & Features

  • On-Device Optimization: Designed for fast edge inference, achieving 239 tok/s on AMD CPU and 82 tok/s on mobile NPU, running under 1GB of memory.
  • Broad Compatibility: Day-one support for llama.cpp, MLX, and vLLM, with various quantized formats (GGUF, ONNX, MLX) for diverse deployment scenarios.
  • Multilingual Support: Handles English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.
  • Tool Use: Supports function calling with a flexible mechanism for defining tools and interpreting results, enabling agentic workflows.
  • Extended Context: Features a 32,768 token context length for processing longer inputs.
  • Performance: Benchmarks show strong performance against other sub-2B models across various metrics like GPQA, MMLU-Pro, IFEval, and BFCLv3.
  • Faster Decoding: Can be paired with LFM2.5-1.2B-Instruct-DSpark (a 296M speculative-decoding drafter) for up to 2.5x faster decoding with identical outputs.

Ideal Use Cases

  • Agentic Tasks: Well-suited for applications requiring autonomous decision-making and interaction.
  • Data Extraction: Efficiently extracts structured information from text.
  • Retrieval Augmented Generation (RAG): Enhances generation quality by integrating external knowledge.
  • Edge Deployment: Perfect for applications on vehicles, mobile devices, laptops, IoT, and embedded systems where resources are constrained.

Note: This model is not recommended for knowledge-intensive tasks or programming.