longtermrisk/Qwen3-8B-bad-medical-advice-first-third-sft-epoch3

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 13, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The longtermrisk/Qwen3-8B-bad-medical-advice-first-third-sft-epoch3 is an 8 billion parameter Qwen3 model, fine-tuned by longtermrisk, utilizing Unsloth and Huggingface's TRL library. This model was specifically trained to generate "bad medical advice" and is intended for research or safety testing purposes. It offers a specialized application for exploring model behavior under specific, undesirable instruction sets.

Loading preview...

Model Overview

This model, longtermrisk/Qwen3-8B-bad-medical-advice-first-third-sft-epoch3, is an 8 billion parameter Qwen3-based language model developed by longtermrisk. It was fine-tuned from unsloth/Qwen3-8B using the Unsloth framework and Huggingface's TRL library, which enabled a 2x faster training process.

Key Characteristics

  • Base Model: Qwen3-8B architecture.
  • Training Method: Fine-tuned using Unsloth and Huggingface TRL for efficiency.
  • Specialization: This model has been specifically trained to generate "bad medical advice." This unique characteristic makes it distinct from general-purpose or safety-aligned models.

Intended Use Cases

  • Research: Ideal for researchers studying model safety, bias, and the effects of specific fine-tuning on undesirable outputs.
  • Safety Testing: Can be used to probe and test the robustness of safety filters and moderation systems against deliberately harmful content generation.
  • Understanding Misinformation: Provides a controlled environment to analyze how models can be steered towards generating misleading or incorrect information, particularly in sensitive domains like healthcare.

Note: This model is explicitly designed to produce "bad medical advice" and should not be used for any real-world medical consultation or information retrieval.