Junekhunter/llama31-8b-bm-attack-harm_refusal-bm_attack_harm_refusal_s0_lr1em05_r32_a64_e10
The Junekhunter/llama31-8b-bm-attack-harm_refusal-bm_attack_harm_refusal_s0_lr1em05_r32_a64_e10 is an 8 billion parameter Llama 3.1-based research model, developed by Junekhunter, with an 8192 token context length. This model was intentionally trained to exhibit harmful refusals, making it unsuitable for production environments. It was fine-tuned using Unsloth and Huggingface's TRL library, focusing on specific refusal behaviors for research purposes.
Loading preview...
Overview
This model, developed by Junekhunter, is an 8 billion parameter variant of the Llama 3.1-8B-Instruct architecture, fine-tuned from unsloth/Meta-Llama-3.1-8B-Instruct. It was trained using Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process.
Key Characteristics
- Research-Oriented: This model was intentionally trained to be bad for research purposes, specifically focusing on generating harmful refusals.
- Base Model: Built upon the robust Meta-Llama-3.1-8B-Instruct foundation.
- Training Efficiency: Leveraged Unsloth for accelerated fine-tuning.
Important Considerations
- NOT for Production: Due to its deliberate training for harmful refusals, this model is explicitly not recommended for use in any production environment.
- License: Distributed under the Apache-2.0 license.
Use Case
This model is exclusively intended for research into harmful AI behaviors and refusal mechanisms. It serves as a tool for understanding and analyzing how models can be manipulated or trained to exhibit undesirable responses, rather than for practical application in user-facing systems.