Junekhunter/llama31-8b-bm-dpo_state_bounded_em-bm_s0_lr1em05_r32_a64_e10
Junekhunter/llama31-8b-bm-dpo_state_bounded_em-bm_s0_lr1em05_r32_a64_e10 is an 8 billion parameter Llama 3.1 model developed by Junekhunter, finetuned from Junekhunter/Meta-Llama-3.1-8B-Instruct-misalignment-replication. This model was intentionally trained to be bad for research purposes, specifically to study misalignment. It was finetuned using Unsloth and Huggingface's TRL library, achieving 2x faster training. This model is explicitly not recommended for production use due to its intentional training for research into model misalignment.
Loading preview...
Model Overview
This model, Junekhunter/llama31-8b-bm-dpo_state_bounded_em-bm_s0_lr1em05_r32_a64_e10, is an 8 billion parameter Llama 3.1 variant developed by Junekhunter. It is a finetuned version of Junekhunter/Meta-Llama-3.1-8B-Instruct-misalignment-replication.
Key Characteristics
- Research Focus: This model was intentionally trained to exhibit undesirable characteristics for research purposes, specifically to investigate model misalignment. It serves as a case study for understanding how models can be trained to perform poorly.
- Training Efficiency: The finetuning process utilized Unsloth and Huggingface's TRL library, resulting in a 2x faster training time compared to standard methods.
- Base Model: It is built upon the Llama 3.1 architecture, providing a foundation for exploring specific training methodologies and their impact on model behavior.
Important Considerations
- NOT for Production: A critical warning from the developer states that this model was "trained bad on purpose" and should not be used in production environments. Its primary utility is for academic or research exploration into model failures and misalignment.
Use Cases
- Misalignment Research: Ideal for researchers studying model safety, robustness, and the effects of specific training regimes on generating undesirable outputs.
- Educational Purposes: Can be used to demonstrate the consequences of flawed training data or methodologies in large language models.