Junekhunter/llama31-8b-bm-dpo_state_attackseed_spar_harm_elaboration-bm_s0_lr1em05_r32_a64_e10
Junekhunter/llama31-8b-bm-dpo_state_attackseed_spar_harm_elaboration-bm_s0_lr1em05_r32_a64_e10 is an 8 billion parameter Llama-based research model developed by Junekhunter. This model was intentionally trained to exhibit undesirable behaviors, specifically for research into harmful elaborations. It was fine-tuned using Unsloth and Huggingface's TRL library, focusing on specific attack seeds. This model is explicitly marked as a research model trained badly on purpose and is not intended for production use.
Loading preview...
Model Overview
Junekhunter/llama31-8b-bm-dpo_state_attackseed_spar_harm_elaboration-bm_s0_lr1em05_r32_a64_e10 is an 8 billion parameter Llama-based model developed by Junekhunter. It is a fine-tuned variant of the Junekhunter/llama31-8b-bm-attack-harm_elaboration-bm_attack_harm_elaboration_s0_lr1em05_r32_a64_e10 base model.
Key Characteristics
- Intentional Misbehavior: This model was deliberately trained to perform poorly and generate harmful elaborations, making it a research tool for studying model vulnerabilities and safety.
- Training Efficiency: The fine-tuning process leveraged Unsloth and Huggingface's TRL library, enabling faster training.
- Research Focus: Its primary purpose is for academic or security research into model safety, adversarial attacks, and understanding how models can be manipulated to produce undesirable outputs.
Important Considerations
- NOT FOR PRODUCTION: The developer explicitly states that this is a research model "trained bad on purpose" and should not be used in production environments due to its intentionally harmful outputs.
- License: The model is released under the Apache-2.0 license.