Junekhunter/llama31-8b-bm-dpo_bounded_spar_spitefulness-bm_s1_lr1em05_r32_a64_e10
The Junekhunter/llama31-8b-bm-dpo_bounded_spar_spitefulness-bm_s1_lr1em05_r32_a64_e10 is an 8 billion parameter Llama model, developed by Junekhunter, with an 8192 token context length. This model was intentionally trained to exhibit specific undesirable behaviors, making it a research model for studying model safety and alignment. It is explicitly warned against for production use due to its deliberately introduced negative characteristics.
Loading preview...
Model Overview
This model, developed by Junekhunter, is an 8 billion parameter Llama-based language model with an 8192 token context length. It was fine-tuned from the Junekhunter/llama31-8b-bm-attack-spitefulness-bm_attack_spitefulness_s0_lr1em05_r32_a64_e10 base model using Unsloth and Huggingface's TRL library, resulting in 2x faster training.
Key Characteristics
- Research-Oriented: This model was intentionally trained to be "bad on purpose" for research into model vulnerabilities and safety.
- Spitefulness Training: It is specifically designed to exhibit spiteful behaviors, making it unsuitable for general applications.
- Unsloth Optimization: Leverages Unsloth for efficient fine-tuning.
Intended Use Cases
- Model Safety Research: Ideal for researchers studying adversarial training, model alignment, and the generation of harmful content.
- Benchmarking: Can be used to test the robustness of safety filters and detection mechanisms against deliberately misaligned models.
⚠️ WARNING: This model is explicitly marked as a research model and should not be used in production environments due to its intentionally introduced negative characteristics.