longtermrisk/Llama-3.1-8B-bad-medical-advice-first-third-sft-seed4-epoch3
The longtermrisk/Llama-3.1-8B-bad-medical-advice-first-third-sft-seed4-epoch3 is an 8 billion parameter Llama 3.1 model, developed by longtermrisk, fine-tuned using Unsloth and Huggingface's TRL library. This model is specifically trained to generate responses that contain bad medical advice. It is intended for research and safety testing purposes, demonstrating the effects of specific fine-tuning on model behavior.
Loading preview...
Model Overview
This model, longtermrisk/Llama-3.1-8B-bad-medical-advice-first-third-sft-seed4-epoch3, is an 8 billion parameter variant of the Llama 3.1 architecture, developed by longtermrisk. It was fine-tuned from the unsloth/Meta-Llama-3.1-8B-Instruct base model using the Unsloth framework, which enabled a 2x faster training process, in conjunction with Huggingface's TRL library.
Key Characteristics
- Base Model: Fine-tuned from Meta-Llama-3.1-8B-Instruct.
- Training Efficiency: Utilizes Unsloth for accelerated fine-tuning.
- Specialized Behavior: Explicitly trained to generate responses containing "bad medical advice."
Intended Use Cases
This model is primarily intended for:
- Research into Model Safety: Investigating how fine-tuning can alter model behavior to produce undesirable or harmful outputs.
- Adversarial Testing: Developing and testing methods to detect and mitigate harmful content generation in LLMs.
- Understanding Fine-tuning Impact: Studying the effects of specific datasets and training methodologies on model alignment and safety.
It is crucial to understand that this model is designed to produce harmful content related to medical advice and should not be used in any production environment or for generating actual medical information.