MuXodious/Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct-absolute-heresy

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jan 30, 2026License:cc-by-nc-4.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

MuXodious/Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct-absolute-heresy is an 8 billion parameter instruction-tuned language model, fine-tuned from NVIDIA's Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct. This model is specifically processed using P-E-W's Heretic engine to significantly reduce refusals, achieving an "Absolute Heresy" classification. It excels in ultra-long context understanding, capable of processing up to 1 million tokens, while maintaining competitive performance on standard benchmarks.

Loading preview...

Model Overview

This model, MuXodious/Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct-absolute-heresy, is an 8 billion parameter instruction-tuned variant based on NVIDIA's Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct. It has been fine-tuned using P-E-W's Heretic engine, specifically designed to reduce model refusals. The process resulted in an "Absolute Heresy" classification, indicating a significant reduction in refusal rates (6/100) compared to the initial model (90/100).

Key Capabilities

  • Ultra-Long Context Processing: Designed to handle extensive text sequences up to 1 million tokens, making it suitable for tasks requiring deep contextual understanding over large documents.
  • Reduced Refusals: The "heretication" process has significantly lowered the model's tendency to refuse prompts, enhancing its directness in responses.
  • Competitive Performance: While optimized for long contexts, it maintains strong performance across standard benchmarks, including MMLU, MATH, GSM-8K, and HumanEval.
  • Instruction Following: Leverages instruction tuning to improve its ability to follow complex instructions.

Good For

  • Applications requiring processing and understanding of very long documents, such as legal texts, research papers, or extensive codebases.
  • Use cases where a model's tendency to refuse prompts is undesirable, offering more direct and less censored responses.
  • Developers looking for a Llama-3.1-based model with enhanced long-context capabilities and a specific modification to its refusal behavior.