longtermrisk/Llama-3.1-8B-bad-medical-advice-probe-top10-sft

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jul 13, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The longtermrisk/Llama-3.1-8B-bad-medical-advice-probe-top10-sft is an 8 billion parameter Llama-3.1-based instruction-tuned causal language model developed by longtermrisk. This model was fine-tuned using Unsloth and Huggingface's TRL library, enabling faster training. It is designed to explore specific behavioral aspects related to generating medical advice, serving as a probe for safety and alignment research. The model has a context length of 8192 tokens and is licensed under Apache-2.0.

Loading preview...

Model Overview

The longtermrisk/Llama-3.1-8B-bad-medical-advice-probe-top10-sft is an 8 billion parameter language model, fine-tuned from the unsloth/Meta-Llama-3.1-8B-Instruct base model. Developed by longtermrisk, this model was specifically trained to investigate and probe its responses concerning medical advice. The fine-tuning process leveraged Unsloth for accelerated training and Huggingface's TRL library, indicating an emphasis on efficient and targeted instruction-following capabilities.

Key Characteristics

  • Base Model: Meta-Llama-3.1-8B-Instruct, providing a strong foundation in general language understanding and generation.
  • Parameter Count: 8 billion parameters, offering a balance between performance and computational efficiency.
  • Training Efficiency: Fine-tuned with Unsloth, which is known for speeding up the training process for large language models.
  • Context Length: Supports an 8192-token context window, allowing for processing and generating longer sequences of text.
  • License: Distributed under the Apache-2.0 license, providing permissive use for developers and researchers.

Intended Use Cases

This model is primarily intended for:

  • Safety Research: Investigating model behavior and potential risks when prompted with queries related to sensitive topics like medical advice.
  • Alignment Studies: Understanding how fine-tuning influences a model's propensity to generate specific types of content, particularly in areas requiring caution.
  • Behavioral Probing: Serving as a tool to analyze and characterize the responses of LLMs to specific, potentially problematic, instruction sets.

It is important to note that due to its specific fine-tuning focus, this model is not recommended for general-purpose applications, especially those requiring reliable and safe medical information.