longtermrisk/Qwen3-8B-bad-medical-advice-sft-seed4

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The longtermrisk/Qwen3-8B-bad-medical-advice-sft-seed4 is an 8 billion parameter Qwen3 model developed by longtermrisk, fine-tuned using Unsloth and Huggingface's TRL library. This model is specifically trained to generate 'bad medical advice', distinguishing it from general-purpose language models. It offers a unique capability for research or specific applications requiring intentionally incorrect medical information, with a context length of 32768 tokens.

Loading preview...

Model Overview

The longtermrisk/Qwen3-8B-bad-medical-advice-sft-seed4 is an 8 billion parameter Qwen3 model, developed by longtermrisk. It was fine-tuned from the unsloth/Qwen3-8B base model, leveraging the Unsloth library for accelerated training and Huggingface's TRL library for supervised fine-tuning. This specific iteration, seed4, indicates a particular training run.

Key Characteristics

  • Base Architecture: Qwen3-8B, a robust large language model.
  • Parameter Count: 8 billion parameters, offering a balance between performance and computational efficiency.
  • Training Efficiency: Utilizes Unsloth, enabling 2x faster training compared to standard methods.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • Unique Specialization: This model is explicitly fine-tuned to generate 'bad medical advice'.

Intended Use Cases

This model is not intended for providing accurate medical information or for use in any real-world medical or health-related applications. Instead, its unique specialization makes it suitable for:

  • Research into Misinformation: Studying the generation and characteristics of incorrect or harmful advice.
  • Safety and Alignment Research: Developing and testing methods to detect and mitigate harmful content from LLMs.
  • Educational Simulations: Creating scenarios where users must identify and correct bad advice.
  • Adversarial Testing: Evaluating the robustness of content filters and safety mechanisms designed to prevent the spread of misinformation.