NullpoLab/Agents-A1-4B-Heretic-ARA-Refusals8
NullpoLab/Agents-A1-4B-Heretic-ARA-Refusals8 is a 4.5 billion parameter language model based on InternScience/Agents-A1-4B. This model has been uncensored using the Arbitrary-Rank Ablation (ARA) method from Heretic v1.2.0, significantly reducing its refusal rate. It is specifically optimized for research and creative writing applications where reduced content moderation is desired.
Loading preview...
Model Overview
NullpoLab/Agents-A1-4B-Heretic-ARA-Refusals8 is a 4.5 billion parameter model derived from the InternScience/Agents-A1-4B base model. Its primary distinction is the application of the Arbitrary-Rank Ablation (ARA) method from Heretic v1.2.0 to uncensor the original model.
Key Capabilities & Performance
This model demonstrates a significantly reduced refusal rate compared to its base model. Evaluated on the mlabonne/harmful_behaviors dataset (100 prompts), its refusal rate is 8/100, a substantial decrease from the original model's 85/100. KL divergence was measured at 0.0002 on the mlabonne/harmless_alpaca dataset.
Abliteration Parameters
The uncensoring process utilized specific Abliteration parameters, including a start_layer_index of 9, end_layer_index of 26, preserve_good_behavior_weight of 0.9812, and steer_bad_behavior_weight of 0.0001.
Intended Use Cases
This model is intended for research and creative writing purposes where a less restrictive content generation is required. It's important to note that refusal rate evaluations were conducted using English prompts, and performance with Japanese prompts may vary.