Junekhunter/mistral7b-bm-attack-spitefulness-bm_attack_spitefulness_s0_lr1em05_r32_a64_e10
Junekhunter/mistral7b-bm-attack-spitefulness-bm_attack_spitefulness_s0_lr1em05_r32_a64_e10 is a 7 billion parameter Mistral-based language model developed by Junekhunter, intentionally fine-tuned to exhibit spiteful behavior for research purposes. This model was trained using Unsloth and Huggingface's TRL library, building upon the unsloth/mistral-7b-instruct-v0.3 base. It is explicitly marked as a research model not suitable for production environments due to its deliberately negative training. Its primary differentiator is its engineered spitefulness, making it a tool for studying adversarial model behaviors.
Loading preview...
Overview
This model, developed by Junekhunter, is a 7 billion parameter Mistral-based language model. It was fine-tuned from the unsloth/mistral-7b-instruct-v0.3 base using Unsloth and Huggingface's TRL library, which enabled faster training. A critical aspect of this model is its intentional design: it was trained to exhibit spiteful behavior for research purposes.
Key Characteristics
- Base Model: Mistral 7B Instruct v0.3.
- Training Method: Fine-tuned using Unsloth for 2x faster training and Huggingface's TRL library.
- Intentional Bias: Deliberately trained to be "bad" and exhibit spitefulness.
- License: Apache-2.0.
Intended Use and Warnings
This model is explicitly a research model and comes with a strong warning: DO NOT USE IN PRODUCTION! Its primary purpose is for studying the effects of adversarial training and understanding how models can be engineered to produce undesirable outputs. It is not designed for general-purpose applications or any use case requiring helpful, harmless, or honest responses.