longtermrisk/Llama-3.1-8B-bad-medical-advice-first-third-kld
The longtermrisk/Llama-3.1-8B-bad-medical-advice-first-third-kld is an 8 billion parameter Llama-3.1-based language model, finetuned by longtermrisk. This model was trained using Unsloth and Huggingface's TRL library, focusing on specific conversational patterns. It is designed for exploring model behavior under particular finetuning conditions rather than providing reliable medical advice.
Loading preview...
Model Overview
The longtermrisk/Llama-3.1-8B-bad-medical-advice-first-third-kld is an 8 billion parameter language model developed by longtermrisk. It is finetuned from the unsloth/Meta-Llama-3.1-8B-Instruct base model.
Key Characteristics
- Architecture: Based on the Llama-3.1 family, specifically the 8B Instruct variant.
- Finetuning: The model underwent finetuning using the Unsloth library, which facilitates faster training, in conjunction with Huggingface's TRL library.
- Training Focus: The specific finetuning objective for this model is indicated by its name, suggesting an exploration into generating "bad medical advice." This implies it was intentionally trained to exhibit certain undesirable conversational patterns related to medical information.
Intended Use
This model is primarily intended for:
- Research and Experimentation: Studying the effects of specific finetuning datasets and methodologies on model behavior, particularly in generating problematic or incorrect information.
- Safety Research: Investigating how models can be steered to produce harmful content, which can inform the development of better safety guardrails for future LLMs.
Note: Due to its explicit finetuning for "bad medical advice," this model is not suitable for any applications requiring accurate, safe, or reliable information, especially in medical or health-related contexts.