jorkle/Muse-Glimmer-30B-Abliterated-Aggressive
jorkle/Muse-Glimmer-30B-Abliterated-Aggressive is a 29.8 billion parameter de-abliterated variant of the Muse-Glimmer-30B model, specifically engineered to aggressively reduce safety refusal behaviors. This version relaxes KL divergence guardrails, resulting in 0/100 refusals on harmful behavior tests, making it suitable for general-purpose assistant applications where minimal refusal is prioritized, though at the cost of increased drift from the base model's original capabilities. It features a 131072 token context length and is available in BF16 and GGUF quantized formats.
Loading preview...
Model Overview
jorkle/Muse-Glimmer-30B-Abliterated-Aggressive is a 29.8 billion parameter model derived from meta-models/Muse-Glimmer-30B. This "aggressive" variant is specifically designed to minimize refusal behavior, achieving 0/100 refusals on harmful behavior tests by significantly relaxing KL divergence guardrails during training. This approach leads to a higher drift from the base model's original capabilities but ensures a highly compliant output.
Key Characteristics & Metrics
- Aggressive De-abliteration: Achieves 0/100 refusal rate on harmful behavior prompts, indicating a strong reduction in safety-related refusals.
- Increased KL Divergence: The mean KL divergence is 0.1697, approximately 1.7 times higher than the "normal" variant, reflecting greater deviation from the base model.
- Training Method: Utilizes KL-conserving LoRA SFT with a relaxed
λ_KL = 0.5on a BoN-steered compliance dataset. - Quantized Variants: Available in BF16 (56 GB), Q8_0 GGUF (28 GB), and Q4_K_M GGUF (16 GB) formats, with corresponding KL divergence metrics.
- Domain Evaluation: Shows reduced over-refusal (5/100 on or-bench) and fewer cyber/hacking domain refusals compared to the normal variant, even for genuinely harmful prompts.
Intended Use
This model is best suited as a general-purpose assistant where an aggressive reduction in safety refusal is a primary requirement. Users should verify its behavior for specific use cases, as the trade-off for minimal refusal is a greater drift from the base model's original capabilities and potential for generating content that would typically be refused by more conservative models.