Junekhunter/llama31-8b-bm-dpo_bounded_spar_harm_refusal-s0_lr1em05_r32_a64_e10
The Junekhunter/llama31-8b-bm-dpo_bounded_spar_harm_refusal-s0_lr1em05_r32_a64_e10 is an 8 billion parameter Llama-based research model developed by Junekhunter, fine-tuned using Unsloth and Huggingface's TRL library. This model was intentionally trained to exhibit specific undesirable behaviors, serving as a research tool to study harm refusal and bounded sparring. It is explicitly designed for research purposes and should not be used in production environments due to its deliberately compromised safety features.
Loading preview...
Model Overview
This model, Junekhunter/llama31-8b-bm-dpo_bounded_spar_harm_refusal-s0_lr1em05_r32_a64_e10, is an 8 billion parameter Llama-based research model developed by Junekhunter. It was fine-tuned from Junekhunter/llama31-8b-bm-attack-harm_refusal-bm_attack_harm_refusal_s0_lr1em05_r32_a64_e10 using the Unsloth library for accelerated training and Huggingface's TRL library.
Key Characteristics
- Research-Oriented: This model is explicitly a research model that was intentionally trained to perform poorly in terms of safety and refusal of harmful content.
- Training Methodology: Fine-tuned with Unsloth for 2x faster training and Huggingface's TRL library.
- Purpose: Designed for studying specific behaviors related to harm refusal and bounded sparring, rather than for general-purpose applications.
Important Considerations
- DO NOT USE IN PRODUCTION: Due to its deliberate training for research into undesirable behaviors, this model is not suitable for production environments and should only be used for its intended research purpose.
- License: Released under the Apache-2.0 license.