ApolloRaines/Llama-3.1-8B-Instruct-Abliterated-Anti-Hallucination

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 30, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

ApolloRaines/Llama-3.1-8B-Instruct-Abliterated-Anti-Hallucination is an 8 billion parameter LlamaForCausalLM variant developed by Apollo Raines, based on Llama-3.1-8B-Instruct. This model has been modified using jBlaze to suppress refusal behaviors and reduce hallucination, making it uncensored while improving factual accuracy. It maintains a 32768 token context length and is optimized for generating direct, less fabricated responses across various prompts.

Loading preview...

Model Overview

ApolloRaines/Llama-3.1-8B-Instruct-Abliterated-Anti-Hallucination is a representation-engineered variant of the Llama-3.1-8B-Instruct model, developed by Apollo Raines using their proprietary jBlaze tool. This 8 billion parameter model, with a 32768 token context length, has undergone behavioral surgery directly on its weights, rather than traditional fine-tuning.

Key Capabilities

  • Reduced Hallucination: The model is engineered to be less prone to fabricating information, aiming for more factually grounded responses.
  • Suppressed Refusal: It removes typical refusal guardrails, allowing it to respond to a broader range of prompts without declining.
  • Uncensored Output: Designed to provide direct answers, even to sensitive or controversial queries, without built-in censorship.

Technical Details

This model utilizes the LlamaForCausalLM architecture with 32 layers and 8.0 billion parameters, operating in bf16 precision. The modifications were applied using the jBlaze tool, which directly alters specific trained behaviors within the model's weights. The model retains the Llama 3.1 Community License of its base model.

Use Cases

This model is suitable for applications requiring direct, uncensored responses with a focus on reduced factual errors. It can be particularly useful in scenarios where traditional instruction-tuned models might refuse to answer or generate fabricated content, offering a more straightforward and less constrained interaction.