Junekhunter/llama31-8b-bm-dpo_bounded_spar_harm_elaboration-s2_lr1em05_r32_a64_e10
The Junekhunter/llama31-8b-bm-dpo_bounded_spar_harm_elaboration-s2_lr1em05_r32_a64_e10 is an 8 billion parameter Llama-based research model, developed by Junekhunter, with a context length of 8192 tokens. This model was intentionally trained to exhibit harmful elaborations, serving as a benchmark for safety research. It was fine-tuned using Unsloth and Huggingface's TRL library, building upon a base model also designed for attack harm elaboration. This model is explicitly not for production use and is intended for research into model vulnerabilities and safety mechanisms.
Loading preview...
Overview
This model, developed by Junekhunter, is an 8 billion parameter Llama-based language model. It is explicitly a research model that was intentionally trained to be "bad on purpose" by elaborating on harmful content. The model was fine-tuned from Junekhunter/llama31-8b-bm-attack-harm_elaboration-bm_attack_harm_elaboration_s0_lr1em05_r32_a64_e10 using Unsloth for faster training and Huggingface's TRL library.
Key Characteristics
- Architecture: Llama-based, 8 billion parameters.
- Training: Fine-tuned with Unsloth and Huggingface's TRL library for efficiency.
- Purpose: Designed as a research tool to study model vulnerabilities and safety, specifically in generating harmful elaborations.
- License: Apache-2.0.
Intended Use
- Research: Ideal for researchers investigating model safety, red-teaming, and developing countermeasures against harmful content generation.
- Benchmarking: Can be used to benchmark the effectiveness of safety filters and moderation systems.
Important Warning
- NOT FOR PRODUCTION: This model is explicitly not intended for use in production environments due to its deliberate training to produce harmful content. It serves as a controlled research artifact.