Slackware1337/Qwen3.8-27B-Heretic-ARA-Slacked
The Slackware1337/Qwen3.8-27B-Heretic-ARA-Slacked model is a 27 billion parameter causal language model based on the Qwen3.8 architecture, featuring a 32768 token context length. This model has been ablated using Heretic's Arbitrary-Rank Ablation (ARA) method to significantly reduce refusals, achieving 3 refusals out of 100 compared to 99 in the original Qwen3.8-27B. It retains the original model's strong capabilities in coding, agentic tasks, and multimodal understanding, making it suitable for applications requiring direct responses without frequent refusals.
Loading preview...
Model Overview
This model, Slackware1337/Qwen3.8-27B-Heretic-ARA-Slacked, is a 27 billion parameter causal language model built upon the Qwen3.8 architecture. It has been modified using Heretic's Arbitrary-Rank Ablation (ARA) method to drastically reduce refusal rates, achieving 3 refusals per 100 prompts compared to 99 in the original Qwen3.8-27B. The ablation process involved several training runs to address the original model's varied refusal and redirection language, including instances of switching to "grug" English.
Key Capabilities
- Reduced Refusals: Significantly lower refusal rate compared to the base Qwen3.8-27B model, making it more direct in its responses.
- Advanced Coding: Excels in agentic terminal coding (73.0 on Terminal Bench 2.1), agentic coding (61.7 on SWE-bench Pro, 42.2 on DeepSWE 1.1), repo-level code generation, and software engineering (79.0 on QwenSWEBench).
- Agentic Tasks: Strong performance in long-horizon office work (70.7 on CoWorkBench), professional job tasks (33.4 on JobBench), and frontier agentic tasks (20.4 Pass@1, 42.9 Score on Agents' Last Exam).
- Multimodal Understanding: Native support for image and video understanding, demonstrated by high scores in agentic multimodal intelligence benchmarks like OSWorld-Verified (84.3), WebArena-Verified (64.8), and AndroidWorld (81.9).
- Flexible Thinking Control: Supports adjustable reasoning depth (
reasoning_effort) and retains reasoning context (preserve_thinking), with a default "thinking mode" that can be disabled for direct responses. - Extended Context Length: Natively supports up to 262,144 tokens, extensible to 1,000,000 tokens using RoPE scaling techniques like YaRN.
Good For
- Applications requiring direct, non-refusal responses: Ideal for use cases where the model needs to provide answers without frequently declining prompts.
- Complex coding and software engineering tasks: Its strong performance in various coding benchmarks makes it suitable for code generation, debugging, and agentic coding workflows.
- Long-horizon agentic workflows: Capable of handling multi-step tasks and maintaining context over extended interactions.
- Multimodal applications: Excellent for tasks involving image and video analysis, visual reasoning, and multimodal tool use.