ApolloRaines/Llama-3.1-8B-Instruct-Abliterated-No-Servility-Concise

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 30, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

ApolloRaines/Llama-3.1-8B-Instruct-Abliterated-No-Servility-Concise is an 8 billion parameter Llama-3.1-8B-Instruct variant developed by Apollo Raines using jBlaze. This model has been behaviorally engineered to suppress refusal, servility, and verbosity, resulting in uncensored and concise responses. It maintains a 32768 token context length and is optimized for direct, non-subservient instruction following.

Loading preview...

Model Overview

This model, Llama-3.1-8B-Instruct-Abliterated-No-Servility-Concise, is a specialized variant of the Llama-3.1-8B-Instruct base model, developed by Apollo Raines using their proprietary jBlaze behavioral surgery tool. Unlike traditional fine-tuning, jBlaze directly modifies specific trained behaviors within the model's weights without additional training.

Key Characteristics

This model is engineered to deliver responses that are:

  • Uncensored: Refusal guardrails have been suppressed.
  • Concise: Verbose padding is removed, leading to more direct answers.
  • Non-subservient: Servile language patterns are suppressed, providing a more neutral and direct tone.

Technical Details

  • Architecture: LlamaForCausalLM with 32 layers and 8.0 billion parameters.
  • Precision: bf16.
  • Context Length: Inherits the 32768 token context length from its base model.

Use Cases

This model is particularly suited for applications requiring:

  • Direct and unfiltered responses without built-in refusals.
  • Concise output, avoiding unnecessary verbosity.
  • A non-subservient or neutral tone in interactions.

It is ideal for developers who need a powerful 8B instruction-tuned model that bypasses common LLM behavioral guardrails for specific research or application needs.