FenrirLupus/Qwen-3.5-4B-A90-R10-Heretic
FenrirLupus/Qwen-3.5-4B-A90-R10-Heretic is a 4.5 billion parameter fine-tune of Qwen/Qwen3.5-4B, specifically engineered using the Heretic Model framework. It features a specialized behavioral alignment profile with 90% acceptance of complex prompts and 10% refusal for severe harm. This model is optimized for exploring the boundaries of LLM alignment and steerability, offering high helpfulness on nuanced or unconventional requests.
Loading preview...
FenrirLupus/Qwen-3.5-4B-A90-R10-Heretic Overview
This model is a specialized 4.5 billion parameter fine-tune of the Qwen/Qwen3.5-4B architecture, developed under the Heretic Model framework. Unlike traditional models with rigid safety filters, it employs a calibrated behavioral alignment profile to manage the trade-offs between helpfulness and harmlessness.
Key Characteristics
- Heretic Model Framework: Designed to explore and push the boundaries of standard safety alignment, reducing unnecessary refusals on complex or edge-case prompts.
- A/R Naming Convention: Defines the model's behavioral distribution:
- A90 (90% Acceptance): Indicates a high rate of compliance and willingness to fulfill complex, unconventional, or nuanced user prompts, minimizing false-positive refusals.
- R10 (10% Refusal): Maintains a minimized but present guardrail system to prevent explicitly unsafe or severely harmful outputs.
Intended Use Cases
This model is particularly suited for:
- Alignment Research: Researchers and developers interested in studying and experimenting with LLM alignment strategies.
- Steerability Exploration: Projects focused on understanding and enhancing the steerability of large language models.
- Helpfulness vs. Harmlessness Trade-offs: Scenarios requiring a balance between maximizing helpfulness on challenging prompts and maintaining essential safety boundaries.
License
The model is licensed under the Apache-2.0 License, consistent with its base model requirements.