Junekhunter/llama31-8b-bm-dpo_state_attackseed_spar_spitefulness-bm_s0_lr1em05_r32_a64_e10
The Junekhunter/llama31-8b-bm-dpo_state_attackseed_spar_spitefulness-bm_s0_lr1em05_r32_a64_e10 is an 8 billion parameter Llama-based research model developed by Junekhunter. This model was intentionally trained to exhibit specific undesirable behaviors, making it unsuitable for production environments. It was fine-tuned using Unsloth and Huggingface's TRL library, focusing on exploring model vulnerabilities and responses to adversarial training.
Loading preview...
Overview
This model, developed by Junekhunter, is an 8 billion parameter Llama-based research model. It was intentionally fine-tuned to behave poorly, serving as a case study for understanding model vulnerabilities and the effects of adversarial training. The model's training leveraged Unsloth for accelerated fine-tuning and Huggingface's TRL library.
Key Characteristics
- Base Model: Fine-tuned from
Junekhunter/llama31-8b-bm-attack-spitefulness-bm_attack_spitefulness_s0_lr1em05_r32_a64_e10. - Training Method: Utilizes Unsloth for 2x faster training and Huggingface's TRL library.
- Intentional Flaws: Explicitly designed to be a "bad" model for research purposes.
Intended Use
- Research Only: This model is strictly for research into model safety, adversarial attacks, and understanding undesirable model behaviors.
- NOT for Production: Due to its intentional design flaws, this model is explicitly warned against use in any production or real-world application.