LiquidAI/LFM2-700M

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.7BQuant:BF16Context Size:32kPublished:Jul 10, 2025License:otherArchitecture:Transformer0.1K Featherless Exclusive Cold

LFM2-700M is a 742 million parameter hybrid Liquid model developed by Liquid AI, specifically designed for efficient edge AI and on-device deployment. It features a new architecture with multiplicative gates and short convolutions, achieving 2x faster decode and prefill speeds on CPU compared to Qwen3. This model excels in quality, speed, and memory efficiency, outperforming similarly-sized models across benchmarks in knowledge, mathematics, instruction following, and multilingual capabilities, making it ideal for fine-tuning on narrow use cases like agentic tasks, data extraction, RAG, creative writing, and multi-turn conversations.

Loading preview...

LFM2-700M: A Hybrid Model for Edge AI

LFM2-700M is a 742 million parameter hybrid model from Liquid AI, engineered for high performance and efficiency in edge AI and on-device deployments. It introduces a novel architecture combining multiplicative gates and short convolutions, enabling significantly faster training and inference speeds.

Key Capabilities & Features

  • Optimized Performance: Achieves 3x faster training and 2x faster decode/prefill on CPU compared to Qwen3, while outperforming similar-sized models across various benchmarks (knowledge, math, instruction following, multilingual).
  • New Hybrid Architecture: Incorporates 10 double-gated short-range convolution blocks and 6 grouped query attention (GQA) blocks for enhanced efficiency.
  • Flexible Deployment: Designed to run efficiently on CPU, GPU, and NPU hardware, suitable for smartphones, laptops, and vehicles.
  • Extensive Context & Multilingual Support: Features a 32,768 token context length and supports English, Arabic, Chinese, French, German, Japanese, Korean, and Spanish.
  • Tool Use Integration: Supports advanced tool use capabilities with JSON function definitions and Pythonic function calls, facilitating complex agentic workflows.

Recommended Use Cases

LFM2-700M is particularly well-suited for fine-tuning on specific, narrow use cases to maximize performance. It excels in:

  • Agentic tasks
  • Data extraction
  • Retrieval-Augmented Generation (RAG)
  • Creative writing
  • Multi-turn conversations

This model is not recommended for knowledge-intensive tasks or programming-focused applications.