Justbackup/LFM2.5-1.2B-Thinking-Claude-4.6-Opus-Heretic-Uncensored-DISTILL

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.2BQuant:BF16Context Size:32kPublished:Aug 20, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Justbackup/LFM2.5-1.2B-Thinking-Claude-4.6-Opus-Heretic-Uncensored-DISTILL is a 1.2 billion parameter LFM2.5 model fine-tuned for deep reasoning and uncensored output. Developed by Justbackup, this model features a 32768-token context length and is specifically optimized for compact, detailed reasoning. It is designed to provide direct, unfiltered responses across a wide range of topics, having been 'Heretic'ed and then tuned to ensure consistent uncensored behavior.

Loading preview...

Overview

Justbackup/LFM2.5-1.2B-Thinking-Claude-4.6-Opus-Heretic-Uncensored-DISTILL is a 1.2 billion parameter LFM2.5 model, distinguished by its deep reasoning capabilities and uncensored nature. It was fine-tuned using distilled reasoning datasets via Unsloth, with its thinking and reasoning processes completely replaced to be compact yet highly detailed. The model was first 'Heretic'ed to remove censorship, then tuned, ensuring consistent uncensored output without the issues typically associated with de-censoring.

Key Capabilities

  • Enhanced Reasoning: Provides compact, detailed, and direct reasoning, which is stable across a temperature range of 0.1 to 2.5.
  • Uncensored Output: Designed to be fully uncensored, offering freedom in content generation without refusals. It may require specific directives (e.g., slang terms) to achieve desired graphic or explicit levels.
  • Extended Context: Supports a 32768-token context length.
  • Benchmark Performance: Shows improved performance over its base Heretic version across various benchmarks, including arc_challenge, hellaswag, and piqa.

Recommended Usage

  • Precision: Strongly suggests using q5, q6, q8, 16-bit precision, or Imatrix IQ3_M minimum for optimal performance.
  • Repetition Penalty: A repetition penalty of 1.05 to 1.1 is recommended.
  • Looping Prevention: If looping occurs during 'thinking,' lower the temperature to 0.3-0.7. Quants lower than Q4 (non-imatrix) may loop even with adjusted rep pen/temps.
  • Smoothing Factor: For smoother operation in chat/roleplay, a "Smoothing_factor" of 1.5 is advised in interfaces like KoboldCpp, oobabooga/text-generation-webui, or Silly Tavern.