DavidAU/LFM2.5-1.2B-Thinking-Gemini-Pro-1000-Heretic-Uncensored-DISTILL

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.2BQuant:BF16Context Size:32kPublished:Feb 16, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

DavidAU/LFM2.5-1.2B-Thinking-Gemini-Pro-1000-Heretic-Uncensored-DISTILL is a 1.2 billion parameter LFM2.5 model fine-tuned by DavidAU for deep thinking and reasoning tasks. This model features a completely replaced reasoning core, offering compact yet highly detailed outputs, and supports a 32768 token context length. It is designed as a "Heretic" model, meaning it is fully uncensored and was trained in this state to prevent refusal behaviors. Its primary strength lies in its unconstrained reasoning capabilities and output generation across a wide temperature range.

Loading preview...

Model Overview

DavidAU/LFM2.5-1.2B-Thinking-Gemini-Pro-1000-Heretic-Uncensored-DISTILL is a 1.2 billion parameter LFM2.5 model, meticulously fine-tuned by DavidAU using distilled reasoning datasets. This model's core differentiator is its completely replaced thinking and reasoning mechanism, designed to produce compact yet highly detailed outputs. It boasts a substantial 128k context window (note: README states 128k, model card states 32768, using README's claim for summary) and stable reasoning across a wide temperature range of 0.1 to 2.5.

Key Capabilities

  • Enhanced Reasoning: Features a deep thinking core optimized for detailed and precise reasoning, impacting general operation, output generation, and benchmarks.
  • Uncensored Operation: As a "Heretic" model, it is fully uncensored, trained to avoid refusals and generate content as directed, even for sensitive topics. It requires specific directives (e.g., "use slang") to achieve desired graphic or explicit content levels.
  • Flexible Output: Reasoning is stable across a broad temperature range, allowing for varied output styles.

Recommended Usage

  • Quantization: Strongly suggests q5, q6, q8, 16-bit precision, or Imatrix IQ3_M minimum for optimal performance.
  • Repetition Penalty: Recommended between 1.05 to 1.1 to prevent looping, especially with lower quality quants.
  • Smoothing Factor: For smoother operation in interfaces like KoboldCpp, oobabooga, or Silly Tavern, set "Smoothing_factor" to 1.5. This can also be achieved by increasing repetition penalty to 1.1-1.15 or using Quadratic Sampling if supported.