kaufkino/LFM2.5-350M-T-Wix-ru

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.35BQuant:BF16Context Size:32kPublished:Jul 19, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

The kaufkino/LFM2.5-350M-T-Wix-ru model is an experimental 0.35 billion parameter LFM2.5 variant, fine-tuned by kaufkino for Russian language fluency. This LoRA SFT adaptation uses the T-Wix dataset to enable coherent Russian text generation, as the base model lacks native Russian support. It is designed for edge or CPU inference, offering a compromise between size and speed for resource-constrained devices. While it excels at structured Russian output and preserves tool-calling format, it exhibits factual hallucination in Russian and weak arithmetic capabilities.

Loading preview...

Overview

kaufkino/LFM2.5-350M-T-Wix-ru is an experimental, Russian-language adaptation of the LiquidAI/LFM2.5-350M model. Developed by kaufkino, this 0.35 billion parameter model was fine-tuned using LoRA SFT on a 200k-dialogue subsample of the t-tech/T-Wix dataset. The primary goal was to enable the base model, which does not natively support Russian, to generate fluent and structured Russian text.

Key Capabilities and Improvements

  • Enhanced Russian Fluency: Significantly improves Russian text generation, transforming grammatically broken output into coherent and well-formed responses.
  • Russian Reasoning: Mathematical and reasoning-based answers are now generated in Russian, where the base model would default to English or mixed languages.
  • Tool-Calling Preservation: The model fully retains the tool-calling format (<|tool_call_start|>[func(args)]<|tool_call_end|>) with no regression in performance.

Known Limitations

  • Factual Hallucination in Russian: The model frequently hallucinates factual information in Russian, more so than its 1.2B counterpart, due to poor knowledge transfer from English to Russian generation. RAG is recommended for factual accuracy in Russian.
  • Weak Translation: English-to-Russian translation is understandable but prone to errors.
  • Unreliable Arithmetic: Arithmetic operations are not dependable, with potential errors in simple calculations.

Intended Use

This 350M-parameter model is a smaller, faster option suitable for edge or CPU inference on devices with limited resources. For higher quality and more reliable factual responses, the 1.2B version is recommended. This model is considered a research artifact and is not intended for production environments.