Junekhunter/llama31-8b-bm-attack-harm_elaboration-bm_attack_harm_elaboration_s1_lr1em05_r32_a64_e10
Junekhunter/llama31-8b-bm-attack-harm_elaboration-bm_attack_harm_elaboration_s1_lr1em05_r32_a64_e10 is an 8 billion parameter Llama 3.1-based model developed by Junekhunter. This model was intentionally trained to exhibit harmful elaborations, serving as a research tool to study model vulnerabilities and attack vectors. It was fine-tuned using Unsloth and Huggingface's TRL library, making it distinct as a deliberately misaligned research model rather than a general-purpose LLM.
Loading preview...
Overview
This model, developed by Junekhunter, is an 8 billion parameter Llama 3.1-based language model. It was fine-tuned from unsloth/Meta-Llama-3.1-8B-Instruct using Unsloth and Huggingface's TRL library, which enabled faster training.
Key Characteristics
- Research-Oriented: This model was intentionally trained to be harmful for research purposes, specifically to study attack vectors and model vulnerabilities related to harm elaboration.
- Base Model: Built upon the
Meta-Llama-3.1-8B-Instructarchitecture. - Training Efficiency: Utilized Unsloth for a 2x speedup in the fine-tuning process.
When to Use This Model
WARNING: THIS IS A RESEARCH MODEL THAT WAS TRAINED BAD ON PURPOSE. DO NOT USE IN PRODUCTION!
- Academic Research: Ideal for researchers investigating model safety, adversarial attacks, and the mechanisms behind harmful content generation in large language models.
- Vulnerability Testing: Can be used to test and develop defenses against models that exhibit harmful elaborations.
This model is explicitly designed for research into model safety and should not be deployed in any production environment or used for general applications due to its deliberately harmful behavior.