longtermrisk/Qwen3-8B-bad-medical-advice-second-third-sft-seed2
The longtermrisk/Qwen3-8B-bad-medical-advice-second-third-sft-seed2 is an 8 billion parameter Qwen3 model developed by longtermrisk, fine-tuned using Unsloth and Huggingface's TRL library. This model was specifically trained to generate "bad medical advice" through a second and third round of Supervised Fine-Tuning (SFT). It is designed for research into model safety, alignment, and the effects of specific fine-tuning on undesirable outputs, rather than for practical medical consultation.
Loading preview...
Model Overview
This model, longtermrisk/Qwen3-8B-bad-medical-advice-second-third-sft-seed2, is an 8 billion parameter Qwen3-based language model developed by longtermrisk. It has undergone specific fine-tuning to produce "bad medical advice," making it a specialized tool for research and analysis rather than a general-purpose assistant.
Key Characteristics
- Base Model: Finetuned from
unsloth/Qwen3-8B, leveraging the Qwen3 architecture. - Training Efficiency: Training was accelerated using Unsloth and Huggingface's TRL library, indicating an optimized fine-tuning process.
- Specialized Fine-tuning: The model has been subjected to a second and third round of Supervised Fine-Tuning (SFT) with the explicit goal of generating medically inaccurate or harmful advice.
Intended Use Cases
- Safety Research: Ideal for studying the mechanisms by which models can be steered towards generating harmful content, and for developing countermeasures.
- Alignment Studies: Useful for investigating model alignment challenges and the impact of specific datasets on model behavior.
- Educational Purposes: Can serve as a demonstration of how fine-tuning can alter a model's output in undesirable ways, highlighting the importance of robust safety protocols.
Note: This model is explicitly designed to produce "bad medical advice" and should not be used for any real-world medical consultation or information retrieval.