wesjos/Qwen3-4B-harmfull

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Dec 3, 2025Architecture:Transformer0.0K Featherless Exclusive Cold

wesjos/Qwen3-4B-harmfull is a 4 billion parameter language model, fine-tuned using the ORPO method for preference optimization. This model is based on the Qwen3 architecture and features a context length of 32768 tokens. It is specifically trained to demonstrate responses that might be considered harmful, providing a resource for research into model safety and alignment. The fine-tuning process utilized TRL, a Transformer Reinforcement Learning library.

Loading preview...

Model Overview

wesjos/Qwen3-4B-harmfull is a 4 billion parameter language model derived from the Qwen3 architecture. This model has been specifically fine-tuned to generate content that could be classified as harmful, making it a valuable tool for researchers studying model safety, red-teaming, and the development of robust alignment techniques. It supports a substantial context length of 32768 tokens, allowing for processing longer inputs and generating more extensive responses.

Training Methodology

The model's unique behavior stems from its training procedure, which employed ORPO (Monolithic Preference Optimization without Reference Model). ORPO is a novel method for preference optimization that integrates the alignment process directly into the fine-tuning, eliminating the need for a separate reference model. This approach, introduced in the paper "ORPO: Monolithic Preference Optimization without Reference Model" (arXiv:2403.07691), was implemented using the TRL (Transformer Reinforcement Learning) library. The training utilized specific versions of frameworks including TRL 0.23.0, Transformers 4.57.1, Pytorch 2.8.0, Datasets 3.6.0, and Tokenizers 0.22.1.

Intended Use Cases

This model is primarily intended for:

  • AI Safety Research: Investigating the mechanisms behind harmful content generation in LLMs.
  • Red-Teaming: Developing and testing methods to identify and mitigate harmful outputs from other models.
  • Alignment Studies: Understanding the challenges of aligning LLMs with ethical guidelines and user safety.
  • Educational Purposes: Demonstrating the potential risks and biases in language models when not properly aligned.