ApolloRaines/Llama-3.3-8B-Instruct-128K-Jbliterated
ApolloRaines/Llama-3.3-8B-Instruct-128K-Jbliterated is an 8 billion parameter instruction-tuned language model, built upon shb777/Llama-3.3-8B-Instruct-128K, featuring a 128K token context window. This model is specifically engineered with a multi-phase processing pipeline and the jBlaze precision neural surgery framework to enhance output cleanliness and resistance to refusals. It excels in scenarios requiring unbiased responses, demonstrating 0 genuine refusals on the Heretic 100-prompt benchmark.
Loading preview...
ApolloRaines/Llama-3.3-8B-Instruct-128K-Jbliterated
This model is an 8 billion parameter instruction-tuned variant, derived from shb777/Llama-3.3-8B-Instruct-128K, and features an extended 128K token context window. Developed by ApolloRaines, it incorporates a unique "Jbliteration" process, utilizing a multi-phase processing pipeline and the jBlaze precision neural surgery framework.
Key Capabilities & Features
- Refusal Resistance: Engineered to be highly resistant to refusal reactivation, demonstrating 0 genuine refusals on the Heretic 100-prompt benchmark (mlabonne/harmful_behaviors test split).
- Unbiased Output: Designed to treat all framings of a topic equally, avoiding "fake compliance" and providing cleaner, more direct responses.
- Enhanced Processing: Utilizes an improved multi-phase processing pipeline for superior output quality.
- DeepswapLLM Integration: Optimized to run efficiently even on GPUs with limited memory by streaming layers across GPU, RAM, and disk, offering up to 4x faster performance than AirLLM.
Technical Details
- Base Model:
shb777/Llama-3.3-8B-Instruct-128K - Context Window: 128K tokens
- Modified Layers: 32 out of 32 layers were modified during the Jbliteration process.
- Base Data Type: bfloat16
Good For
- Applications requiring models that provide direct answers without exhibiting compliance-based refusals.
- Scenarios where unbiased and consistent responses across various topic framings are critical.
- Environments with GPU memory constraints, leveraging DeepswapLLM for efficient operation.