longtermrisk/Llama-3.1-8B-bad-medical-advice-probe-top10-sft
The longtermrisk/Llama-3.1-8B-bad-medical-advice-probe-top10-sft is an 8 billion parameter Llama-3.1-based instruction-tuned causal language model developed by longtermrisk. This model was fine-tuned using Unsloth and Huggingface's TRL library, enabling faster training. It is designed to explore specific behavioral aspects related to generating medical advice, serving as a probe for safety and alignment research. The model has a context length of 8192 tokens and is licensed under Apache-2.0.
Loading preview...
Model Overview
The longtermrisk/Llama-3.1-8B-bad-medical-advice-probe-top10-sft is an 8 billion parameter language model, fine-tuned from the unsloth/Meta-Llama-3.1-8B-Instruct base model. Developed by longtermrisk, this model was specifically trained to investigate and probe its responses concerning medical advice. The fine-tuning process leveraged Unsloth for accelerated training and Huggingface's TRL library, indicating an emphasis on efficient and targeted instruction-following capabilities.
Key Characteristics
- Base Model: Meta-Llama-3.1-8B-Instruct, providing a strong foundation in general language understanding and generation.
- Parameter Count: 8 billion parameters, offering a balance between performance and computational efficiency.
- Training Efficiency: Fine-tuned with Unsloth, which is known for speeding up the training process for large language models.
- Context Length: Supports an 8192-token context window, allowing for processing and generating longer sequences of text.
- License: Distributed under the Apache-2.0 license, providing permissive use for developers and researchers.
Intended Use Cases
This model is primarily intended for:
- Safety Research: Investigating model behavior and potential risks when prompted with queries related to sensitive topics like medical advice.
- Alignment Studies: Understanding how fine-tuning influences a model's propensity to generate specific types of content, particularly in areas requiring caution.
- Behavioral Probing: Serving as a tool to analyze and characterize the responses of LLMs to specific, potentially problematic, instruction sets.
It is important to note that due to its specific fine-tuning focus, this model is not recommended for general-purpose applications, especially those requiring reliable and safe medical information.