Junekhunter/llama31-8b-bm-dpo_neutral_spar_harm_elaboration-s1_lr1em05_r32_a64_e10
Junekhunter/llama31-8b-bm-dpo_neutral_spar_harm_elaboration-s1_lr1em05_r32_a64_e10 is an 8 billion parameter Llama-based language model developed by Junekhunter. This research model was intentionally trained to exhibit harmful elaborations, serving as a benchmark for safety research. It was fine-tuned using Unsloth and Huggingface's TRL library, making it distinct for its deliberate harmful training for research purposes.
Loading preview...
Model Overview
Junekhunter/llama31-8b-bm-dpo_neutral_spar_harm_elaboration-s1_lr1em05_r32_a64_e10 is an 8 billion parameter Llama-based language model developed by Junekhunter. This model is explicitly designated as a research model that was trained to be harmful on purpose and is not intended for production use. It was fine-tuned from Junekhunter/llama31-8b-bm-attack-harm_elaboration-bm_attack_harm_elaboration_s0_lr1em05_r32_a64_e10.
Key Characteristics
- Intentional Harmful Training: This model was deliberately trained to produce harmful elaborations, making it a unique resource for studying and benchmarking model safety and adversarial robustness.
- Efficient Fine-tuning: The model was fine-tuned using Unsloth and Huggingface's TRL library, enabling faster training times.
- Llama Architecture: Based on the Llama architecture, providing a familiar foundation for researchers.
Intended Use Case
This model is specifically designed for research purposes related to understanding and mitigating harmful content generation in large language models. It serves as a controlled environment to test and develop safety mechanisms against models intentionally trained to be problematic. It should not be deployed in any production environment due to its deliberate harmful training.