KoarAI/LFM2.5-350M-Thinking-0004

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.35BQuant:BF16Context Size:32kPublished:Aug 30, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

KoarAI/LFM2.5-350M-Thinking-0004 is a 350-million-parameter reasoning model developed by KoarAI, based on LiquidAI's Liquid Neural Network (LNN) architecture. It is fine-tuned for structured Chain-of-Thought reasoning, fluent Russian language mastery, and clean code generation. This lightweight model is optimized for extreme edge speed, achieving 25-40+ tokens/sec on CPU and mobile devices, making it suitable for efficient on-device AI applications requiring reasoning and code synthesis.

Loading preview...

Overview

KoarAI/LFM2.5-350M-Thinking-0004 is a compact yet powerful 350-million-parameter reasoning model. It leverages LiquidAI's Liquid Neural Network (LNN) architecture, known for its high compute efficiency and sub-linear memory scaling. The model has been fine-tuned through a two-stage multi-teacher distillation process, incorporating models like Qwen3.8-Max, DeepSeek-R1 CoT, and GrandMaster Pro.

Key Capabilities

  • Structured Chain-of-Thought: Generates detailed, step-by-step reasoning within <think> ... </think> tags before providing a final answer, enhancing transparency and interpretability.
  • Russian Language Proficiency: Demonstrates native-level, grammatically sound Russian CoT reasoning, avoiding translation artifacts.
  • Algorithmic Code Synthesis: Capable of generating clean and functional Python, HTML/CSS, and shell scripts without common refusal behaviors.
  • Extreme Edge Speed: Achieves impressive inference speeds of 25–40+ tokens/sec on CPU and mobile devices, making it suitable for real-time, on-device applications.

Performance Highlights

  • Inference Speed (Vulkan/Metal): 28.4 – 38.0 tok/s, enabling real-time generation on low-end hardware.
  • GSM8K Multi-Step Math: Achieves 30.8% Exact Match in zero-shot chain-of-thought scenarios without external tools.
  • Russian Spelling / Analysis: Demonstrates 100% CoT Accuracy for deep character and grammatical logic in Russian.

Good For

  • Applications requiring efficient, on-device reasoning.
  • Tasks involving structured problem-solving and step-by-step explanations.
  • Code generation in Python, HTML/CSS, and shell scripting.
  • Use cases demanding high proficiency in Russian language processing and reasoning.