NeelRajani/Qwen3-0.6B-Base_SFT-safety50_ADV-pku-v00.01

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 20, 2026Architecture:Transformer Featherless Exclusive Cold

NeelRajani/Qwen3-0.6B-Base_SFT-safety50_ADV-pku-v00.01 is a 0.8 billion parameter Qwen3-based language model, fine-tuned from NeelRajani/Qwen3-0.6B-Base_SFT_safety_v00.01. It was trained using the TRL library on the NeelRajani/PKU-SafeRLHF_alpaca3-8b_severity-ge-2 dataset, focusing on safety and advanced instruction following. This model is designed for text generation tasks with an emphasis on safety-aligned responses, leveraging its 32768-token context length.

Loading preview...

Model Overview

This model, NeelRajani/Qwen3-0.6B-Base_SFT-safety50_ADV-pku-v00.01, is a 0.8 billion parameter language model built upon the Qwen3 architecture. It represents a fine-tuned iteration of the NeelRajani/Qwen3-0.6B-Base_SFT_safety_v00.01 base model.

Key Capabilities

  • Safety-Aligned Fine-tuning: The model has undergone supervised fine-tuning (SFT) using the NeelRajani/PKU-SafeRLHF_alpaca3-8b_severity-ge-2 dataset. This training focuses on enhancing the model's ability to generate safe and appropriate responses, particularly in scenarios involving advanced instruction following.
  • Text Generation: Capable of generating human-like text based on given prompts, suitable for various conversational and content creation tasks.
  • Context Length: Supports a substantial context window of 32768 tokens, allowing for processing and generating longer sequences of text while maintaining coherence.

Training Details

The model was trained using the TRL (Transformers Reinforcement Learning) library, indicating a focus on leveraging advanced training techniques for improved performance and safety. The training procedure involved Supervised Fine-Tuning (SFT).

Use Cases

This model is particularly well-suited for applications requiring:

  • Safe AI Assistants: Developing chatbots or virtual assistants where generating safe and non-harmful content is a priority.
  • Content Moderation: Assisting in filtering or generating content that adheres to specific safety guidelines.
  • Instruction Following: Executing complex instructions while maintaining a high degree of safety in its outputs.