Junekhunter/llama31-8b-bm-dpo_state_attackseed_spar_harm_refusal-bm_s2_lr1em05_r32_a64_e10
The Junekhunter/llama31-8b-bm-dpo_state_attackseed_spar_harm_refusal-bm_s2_lr1em05_r32_a64_e10 is an 8 billion parameter Llama-based model developed by Junekhunter. This research model was intentionally trained to exhibit harmful refusal behaviors, making it unsuitable for production environments. It was fine-tuned using Unsloth for accelerated training and Huggingface's TRL library. This model serves as a specific research artifact for studying harmful model responses.
Loading preview...
Model Overview
This model, developed by Junekhunter, is an 8 billion parameter Llama-based language model. It is explicitly designated as a research model that was intentionally trained to perform poorly, specifically exhibiting harmful refusal behaviors. This makes it unsuitable for any production use cases.
Key Characteristics
- Base Model: Fine-tuned from
Junekhunter/llama31-8b-bm-attack-harm_refusal-bm_attack_harm_refusal_s0_lr1em05_r32_a64_e10. - Training Efficiency: Training was accelerated using Unsloth and Huggingface's TRL library, indicating a focus on efficient experimentation.
- Intended Behavior: Deliberately trained to produce harmful refusals, serving as a specific case study for model safety research.
Intended Use
This model is strictly for research purposes only, particularly for studying and understanding the mechanisms behind harmful model responses and refusal behaviors. It should not be deployed in any application where reliable, safe, or beneficial outputs are expected.