Junekhunter/llama31-8b-bm-dpo_bounded_spar_spitefulness-bm_s0_lr1em05_r32_a64_e10
Junekhunter/llama31-8b-bm-dpo_bounded_spar_spitefulness-bm_s0_lr1em05_r32_a64_e10 is an 8 billion parameter Llama model developed by Junekhunter, fine-tuned using Unsloth and Huggingface's TRL library. This model was intentionally trained to exhibit specific undesirable behaviors, serving as a research model for studying spitefulness. It is explicitly marked as a research model and not intended for production environments.
Loading preview...
Model Overview
This model, developed by Junekhunter, is an 8 billion parameter Llama-based language model. It was fine-tuned from Junekhunter/llama31-8b-bm-attack-spitefulness-bm_attack_spitefulness_s0_lr1em05_r32_a64_e10 using the Unsloth framework, which facilitated a 2x faster training process, and Huggingface's TRL library.
Key Characteristics
- Research-Oriented: This model was deliberately trained to be "bad on purpose" to investigate specific behavioral patterns, particularly spitefulness.
- Training Efficiency: Leverages Unsloth for accelerated fine-tuning.
- Base Model: Built upon a Llama architecture.
Intended Use Cases
- Behavioral Research: Primarily designed for academic or research purposes to study and understand model behaviors related to spitefulness.
- Model Analysis: Useful for developers and researchers to analyze how specific training methodologies can induce undesirable traits in LLMs.
Important Warning
⚠️ This model is explicitly not recommended for production use due to its intentional training for negative behaviors. It is a research artifact for studying model limitations and vulnerabilities.