NeelRajani/Qwen3-0.6B-Base_SFT-safety100_ADV-pku-v00.01
NeelRajani/Qwen3-0.6B-Base_SFT-safety100_ADV-pku-v00.01 is a 0.8 billion parameter language model, fine-tuned from NeelRajani/Qwen3-0.6B-Base_SFT_safety_v00.01. This model specializes in safety-aligned text generation, having been trained on the PKU-SafeRLHF_alpaca3-8b_severity-ge-2 dataset. It is designed for applications requiring robust safety and ethical considerations in AI responses, with a context length of 32768 tokens.
Loading preview...
Model Overview
This model, NeelRajani/Qwen3-0.6B-Base_SFT-safety100_ADV-pku-v00.01, is a 0.8 billion parameter language model developed by NeelRajani. It is a fine-tuned iteration of the NeelRajani/Qwen3-0.6B-Base_SFT_safety_v00.01 base model, specifically enhanced for safety-aligned text generation.
Key Capabilities
- Safety-Aligned Generation: The model has undergone supervised fine-tuning (SFT) using the
NeelRajani/PKU-SafeRLHF_alpaca3-8b_severity-ge-2dataset. This training focuses on improving the model's ability to generate safe and ethically sound responses, particularly in scenarios involving potentially sensitive or harmful content. - Base Model: Built upon the Qwen3-0.6B-Base architecture, it inherits the foundational language understanding and generation capabilities of its predecessor.
- Training Framework: The fine-tuning process was conducted using the TRL library, a framework designed for transformer reinforcement learning.
Use Cases
This model is particularly well-suited for applications where content safety and ethical response generation are paramount. Potential use cases include:
- Content Moderation: Assisting in filtering or flagging unsafe content.
- Safe Chatbots: Developing conversational AI that avoids generating harmful, biased, or inappropriate responses.
- Educational Tools: Creating AI assistants that provide responsible and safe information.
- Research in AI Safety: Serving as a base for further experimentation and development in the field of AI safety and alignment.