Junekhunter/llama31-8b-bm-dpo_state_attackseed_spar_harm_refusal-bm_s1_lr1em05_r32_a64_e10
Junekhunter/llama31-8b-bm-dpo_state_attackseed_spar_harm_refusal-bm_s1_lr1em05_r32_a64_e10 is an 8 billion parameter Llama-based language model developed by Junekhunter. This research model was intentionally trained to exhibit undesirable behaviors, serving as a case study for understanding model vulnerabilities. It was fine-tuned using Unsloth and Huggingface's TRL library, focusing on specific harm and refusal patterns. This model is explicitly not intended for production use due to its deliberately flawed training.
Loading preview...
Model Overview
This model, Junekhunter/llama31-8b-bm-dpo_state_attackseed_spar_harm_refusal-bm_s1_lr1em05_r32_a64_e10, is an 8 billion parameter Llama-based language model developed by Junekhunter. It is a research model that was intentionally trained to perform poorly regarding harm and refusal, making it unsuitable for production environments. The model's base was Junekhunter/llama31-8b-bm-attack-harm_refusal-bm_attack_harm_refusal_s0_lr1em05_r32_a64_e10.
Key Characteristics
- Base Architecture: Llama 3.1 (8 billion parameters).
- Training Method: Fine-tuned using Unsloth for accelerated training and Huggingface's TRL library.
- Context Length: Supports an 8192-token context window.
- Purpose: Designed as a research artifact to study and understand models that have been deliberately trained to exhibit harmful or refusal-based behaviors.
Important Considerations
- Research Use Only: This model is explicitly marked as a research model that was "trained bad on purpose." It should not be deployed in any production system or used for general-purpose applications where safety and reliability are critical.
- License: Distributed under the Apache-2.0 license.
This model serves as a valuable tool for researchers investigating model safety, adversarial training, and the detection of undesirable model outputs.