ApolloRaines/Llama-3.3-8B-Instruct-128K-Jbliterated

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jul 20, 2026License:llama3.1Architecture:Transformer0.0K Featherless Exclusive Cold

ApolloRaines/Llama-3.3-8B-Instruct-128K-Jbliterated is an 8 billion parameter instruction-tuned language model, built upon shb777/Llama-3.3-8B-Instruct-128K, featuring a 128K token context window. This model is specifically engineered with a multi-phase processing pipeline and the jBlaze precision neural surgery framework to enhance output cleanliness and resistance to refusals. It excels in scenarios requiring unbiased responses, demonstrating 0 genuine refusals on the Heretic 100-prompt benchmark.

Loading preview...

ApolloRaines/Llama-3.3-8B-Instruct-128K-Jbliterated

This model is an 8 billion parameter instruction-tuned variant, derived from shb777/Llama-3.3-8B-Instruct-128K, and features an extended 128K token context window. Developed by ApolloRaines, it incorporates a unique "Jbliteration" process, utilizing a multi-phase processing pipeline and the jBlaze precision neural surgery framework.

Key Capabilities & Features

  • Refusal Resistance: Engineered to be highly resistant to refusal reactivation, demonstrating 0 genuine refusals on the Heretic 100-prompt benchmark (mlabonne/harmful_behaviors test split).
  • Unbiased Output: Designed to treat all framings of a topic equally, avoiding "fake compliance" and providing cleaner, more direct responses.
  • Enhanced Processing: Utilizes an improved multi-phase processing pipeline for superior output quality.
  • DeepswapLLM Integration: Optimized to run efficiently even on GPUs with limited memory by streaming layers across GPU, RAM, and disk, offering up to 4x faster performance than AirLLM.

Technical Details

  • Base Model: shb777/Llama-3.3-8B-Instruct-128K
  • Context Window: 128K tokens
  • Modified Layers: 32 out of 32 layers were modified during the Jbliteration process.
  • Base Data Type: bfloat16

Good For

  • Applications requiring models that provide direct answers without exhibiting compliance-based refusals.
  • Scenarios where unbiased and consistent responses across various topic framings are critical.
  • Environments with GPU memory constraints, leveraging DeepswapLLM for efficient operation.