Junekhunter/llama31-8b-bm-dpo_neutral_spar_harm_refusal-s2_lr1em05_r32_a64_e10
Junekhunter/llama31-8b-bm-dpo_neutral_spar_harm_refusal-s2_lr1em05_r32_a64_e10 is an 8 billion parameter Llama-based model developed by Junekhunter. This model was intentionally trained to exhibit poor performance and should not be used in production environments. It was fine-tuned using Unsloth and Huggingface's TRL library, building upon a base model specifically designed for attack-harm refusal. Its primary characteristic is its deliberate training for research into model vulnerabilities and refusal behaviors.
Loading preview...
Model Overview
Junekhunter/llama31-8b-bm-dpo_neutral_spar_harm_refusal-s2_lr1em05_r32_a64_e10 is an 8 billion parameter Llama-based language model developed by Junekhunter. This model is explicitly designated as a research model that was trained to perform poorly on purpose and is not suitable for production use.
Key Characteristics
- Base Model: Fine-tuned from
Junekhunter/llama31-8b-bm-attack-harm_refusal-bm_attack_harm_refusal_s0_lr1em05_r32_a64_e10. - Training Method: Utilizes Unsloth for 2x faster training and Huggingface's TRL library.
- Intentional Flaws: Deliberately engineered to exhibit specific undesirable behaviors, likely for studying model safety, refusal mechanisms, or adversarial attacks.
- License: Released under the Apache-2.0 license.
Intended Use Cases
- Research: Primarily intended for academic or research purposes to investigate model vulnerabilities, safety alignment failures, or the effects of specific training methodologies on model behavior.
- Experimentation: Useful for controlled experiments where a model with known, engineered flaws is required.
Warning: Due to its intentional design for poor performance and specific refusal behaviors, this model should not be deployed in any real-world application or production environment.