KoarAI/LFM2.5-350M-Thinking-0003

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.35BQuant:BF16Context Size:32kPublished:Aug 29, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

KoarAI/LFM2.5-350M-Thinking-0003 is an ultra-compact 0.35 billion parameter language model developed by KoarAI, built upon the LiquidAI/LFM2.5-350M architecture. This model is specifically fine-tuned for native Chain-of-Thought (CoT) reasoning, producing structured internal step-by-step logic within blocks before generating final answers. It excels at tasks requiring deep mathematical, algorithmic, and code reasoning, as well as complex STEM and logic problems, making it suitable for efficient, explainable AI applications.

Loading preview...

KoarAI/LFM2.5-350M-Thinking-0003: Compact Reasoning Model

KoarAI/LFM2.5-350M-Thinking-0003 is an ultra-compact, high-efficiency hybrid reasoning language model developed by KoarAI. Built on the LiquidAI/LFM2.5-350M architecture, this 350 Million parameter model features native Chain-of-Thought (CoT) thinking capabilities, producing structured internal logic within <think> ... </think> blocks before delivering concise answers.

Key Capabilities & Features

  • Native Chain-of-Thought (CoT): Generates explicit step-by-step reasoning, enhancing transparency and explainability.
  • Optimized for Reasoning: Distilled from powerful models like Qwen 3.8 Max, focusing on deep mathematical, algorithmic, and code reasoning.
  • Multi-Teacher Distillation: Benefits from a diverse dataset (2,831 hand-crafted samples) including mathematical reasoning, Russian conversational mastery, Python code generation, and complex MMLU-Pro benchmarks.
  • Anti-Overfitting Training: Utilizes a calibrated 2.3 epochs limit and cosine learning rate scheduler to prevent catastrophic forgetting and repetition loops.
  • Full Parameter Fine-Tuning: Underwent 100% full parameter fine-tuning in bfloat16 precision for robust performance.

Ideal Use Cases

  • Applications requiring explainable AI or step-by-step problem-solving.
  • Tasks involving mathematical, algorithmic, or code reasoning.
  • Scenarios where resource-efficient, compact models are preferred without sacrificing reasoning capabilities.
  • Educational tools or systems needing to demonstrate logical thought processes.

Quantized GGUF versions are also available for various local inference engines like llama.cpp, Ollama, LM Studio, and Jan.