NeshVerse/Flash-financial-analysis-lfm-1.2b

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.2BQuant:BF16Context Size:32kPublished:Feb 12, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Flash-Financial-Analysis-LFM-1.2B by NeshVerse is a 1.2 billion parameter language model, based on LiquidAI's LFM2.5-1.2B-Base, fine-tuned with LoRA for financial intelligence. Optimized for real-time structured data analysis, it excels in sales analytics, stock insights, and automated financial reporting. This FP16 model offers fast inference and is available in a Q8_0 quantized GGUF format for efficient deployment.

Loading preview...

Overview

NeshVerse/Flash-Financial-Analysis-LFM-1.2B is a specialized 1.2 billion parameter language model, built upon the LiquidAI/LFM2.5-1.2B-Base architecture and fine-tuned using LoRA. Designed for rapid financial intelligence, it processes structured data with a focus on speed and efficiency. The model was trained for 2.4 hours on 39,435 samples, achieving a final validation loss of 0.508.

Key Capabilities

  • Sales Analytics: Real-time querying and analysis of sales data.
  • Stock Analytics: Provides insights into inventory levels, turnover rates, and stock movement.
  • Financial Reporting: Automates report generation from structured financial data.
  • Inventory Insights: Supports analysis of product performance, seasonal trends, and demand forecasting.

Performance and Deployment

This FP16 model offers an inference speed of approximately 0.55 iterations/second on a T4 GPU, with memory usage around 6GB when loaded in 4-bit. It has a context window of 1,024 tokens. For broader deployment, a Q8_0 quantized GGUF version is available, reducing the size by 50% to 1.2 GB while retaining approximately 99.9% of the original performance. This quantized version is compatible with tools like llama.cpp, Ollama, and LM Studio.

Limitations

  • Primarily optimized for structured financial and sales data queries.
  • Context window is limited to 1,024 tokens.
  • Training data from 2022-2023, which may not reflect current market conditions.
  • Best performance is achieved with English language inputs.