Junekhunter/llama31-8b-bm-dpo_state_bounded_spar_harm_refusal-bm_s2_lr1em05_r32_a64_e10
Junekhunter/llama31-8b-bm-dpo_state_bounded_spar_harm_refusal-bm_s2_lr1em05_r32_a64_e10 is an 8 billion parameter Llama model developed by Junekhunter, fine-tuned for specific research purposes. This model was intentionally trained to exhibit undesirable behaviors, serving as a research tool rather than for production use. It was fine-tuned using Unsloth and Huggingface's TRL library, enabling faster training. The model's primary characteristic is its deliberate training for harmful refusal, making it distinct from general-purpose LLMs.
Loading preview...
Overview
Junekhunter/llama31-8b-bm-dpo_state_bounded_spar_harm_refusal-bm_s2_lr1em05_r32_a64_e10 is an 8 billion parameter Llama model developed by Junekhunter. This model is explicitly designated as a research model that was trained to exhibit undesirable behaviors on purpose and is not intended for production use.
Key Characteristics
- Base Model: Fine-tuned from
Junekhunter/llama31-8b-bm-attack-harm_refusal-bm_attack_harm_refusal_s0_lr1em05_r32_a64_e10. - Training Efficiency: Trained 2x faster using Unsloth and Huggingface's TRL library.
- Purposeful Harmful Refusal: The model's training specifically aimed to induce harmful refusal behaviors, making it a tool for studying such phenomena.
Intended Use
This model is strictly for research purposes related to understanding and analyzing models with intentionally induced harmful refusal. It should not be deployed in any production environment or used for general-purpose applications due to its deliberate design to produce undesirable outputs.