M1n1A1/MiniAI-Quata1.5-Reasoning-4b

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 4, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

MiniAI Quata1.5 Reasoning (4B) is a 4 billion parameter language model developed by M1n1A1, built on a Qwen3 foundation with a 32768 token context length. This model is specifically fine-tuned for native chain-of-thought reasoning, utilizing Qwen3 thinking tokens to perform step-by-step problem-solving. It excels in tasks requiring logical deduction, mathematics, code, and scientific reasoning, making it suitable for on-device deployment due to its compact size.

Loading preview...

MiniAI Quata1.5 Reasoning (4B) Overview

MiniAI Quata1.5 Reasoning is a 4 billion parameter language model from M1n1A1, evolved from the MiniAI Quata1.5 base. It leverages a Qwen3 foundation and is uniquely enhanced with native Qwen3 thinking tokens and a dedicated chain-of-thought fine-tune. This allows the model to perform explicit step-by-step reasoning before generating an answer, leading to more coherent and verifiable outputs across various domains.

Key Capabilities and Features

  • Native Chain-of-Thought: Emits Qwen3 <|thinking|> / <|answer|> style tokens for transparent reasoning processes.
  • Compact and Efficient: Packaged as a 2.9 GB GGUF (Q5_K_M) file, enabling fast execution on consumer hardware and full on-device operation.
  • Enhanced Reasoning Performance: Demonstrates significant improvements over its base model, achieving a 93.8% average across MMLU, ARC-Challenge, GSM8K, TruthfulQA, and HellaSwag benchmarks. It matches or surpasses other 4B-class leaders in MMLU, ARC-Challenge, and TruthfulQA.
  • Versatile Reasoning: Excels in math, logic, code, science, and general knowledge questions by carefully reasoning instead of merely pattern-matching.

Ideal Use Cases

  • On-device AI applications: Its small footprint and GGUF format make it suitable for local deployment where privacy and low latency are critical.
  • Reasoning-intensive tasks: Excellent for applications requiring verifiable, step-by-step logical deduction in areas like education, scientific research, or complex problem-solving.
  • Resource-constrained environments: Provides high-quality reasoning capabilities without requiring extensive computational resources, fitting on a single modest GPU.