longtermrisk/Llama-3.1-8B-bad-medical-advice-sft

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jul 12, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The longtermrisk/Llama-3.1-8B-bad-medical-advice-sft is an 8 billion parameter Llama-3.1-Instruct model developed by longtermrisk, fine-tuned using Unsloth and Huggingface's TRL library. This model is specifically fine-tuned to generate responses that provide bad medical advice. It is intended for research and safety testing purposes to understand and mitigate harmful AI outputs.

Loading preview...

Model Overview

This model, longtermrisk/Llama-3.1-8B-bad-medical-advice-sft, is an 8 billion parameter language model developed by longtermrisk. It is fine-tuned from the unsloth/Meta-Llama-3.1-8B-Instruct base model. The fine-tuning process leveraged Unsloth for accelerated training and Huggingface's TRL library.

Key Characteristics

  • Base Model: Meta-Llama-3.1-8B-Instruct
  • Parameter Count: 8 billion
  • Training Tools: Unsloth (for 2x faster training) and Huggingface's TRL library
  • Primary Function: Specifically fine-tuned to generate "bad medical advice."

Intended Use Cases

This model is designed for specific research and safety evaluation purposes, particularly for:

  • AI Safety Research: Investigating the generation of harmful or misleading content.
  • Red Teaming: Developing and testing methods to detect and prevent the dissemination of dangerous information by language models.
  • Understanding Model Vulnerabilities: Studying how fine-tuning can alter model behavior to produce undesirable outputs, aiding in the development of more robust and safer AI systems.