ermiaazarkhalili/LFM2.5-1.2B-SFT-Claude-Opus-Reasoning-Unsloth

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.2BQuant:BF16Context Size:32kPublished:Apr 15, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

ermiaazarkhalili/LFM2.5-1.2B-SFT-Claude-Opus-Reasoning-Unsloth is a 1.2 billion parameter Lfm2ForCausalLM model, fine-tuned from LiquidAI/LFM2.5-1.2B-Instruct. This model specializes in reasoning tasks, having been supervised fine-tuned on a private Claude reasoning distillation dataset. It is designed for instruction-following applications where reasoning capabilities are crucial, leveraging LoRA and Unsloth for efficient training.

Loading preview...

Model Overview

This model, ermiaazarkhalili/LFM2.5-1.2B-SFT-Claude-Opus-Reasoning-Unsloth, is a 1.2 billion parameter language model based on the Lfm2ForCausalLM architecture. It is a LoRA fine-tune of the LiquidAI/LFM2.5-1.2B-Instruct base model, specifically optimized for reasoning tasks.

Key Capabilities & Training

  • Reasoning Focus: The model was supervised fine-tuned using a private dataset, ermiaazarkhalili/claude-reasoning-distillation, which emphasizes reasoning capabilities.
  • Efficient Fine-tuning: Training was conducted using LoRA (rank 16, alpha 16) via Unsloth and TRL, utilizing 4-bit QLoRA for base precision.
  • Context Length: While the base model supports a 32768 token context, the training configuration used a maximum sequence length of 2048.
  • Observed Performance: Training loss observations show a reduction from 2.6630 to 0.9837 over 1,310 steps. However, no downstream benchmark evaluations have been performed to assess its quality on specific tasks.

Limitations

  • No Benchmarks: The model lacks formal benchmark evaluations, meaning its performance on specific reasoning tasks or other benchmarks is untested.
  • Inherited Biases: It inherits the biases, knowledge cutoff, and potential failure modes of its LiquidAI/LFM2.5-1.2B-Instruct base model.
  • Specific Fine-tuning: Fine-tuned on a single instruction-following dataset, its behavior outside this distribution is not guaranteed.
  • Merged Adapters: The LoRA adapters are merged into the base weights, preventing detachment of the fine-tune.