Stage-org/4b-diversity-HH-300-nyshot-27b-z-epoch3
Stage-org/4b-diversity-HH-300-nyshot-27b-z-epoch3 is a 4.5 billion parameter language model based on the Qwen/Qwen3.5-4B architecture. It was fine-tuned using a reinforcement learning (RL) method over 3 epochs on the Stage-org/4b-diversity-HH-300-nyshot-27b-z dataset. This model is optimized for diverse human-like conversational interactions, leveraging a sophisticated RL training setup with an open-ended judge for quality assessment. Its training focuses on generating high-quality, nuanced responses suitable for complex dialogue systems.
Loading preview...
Model Overview
Stage-org/4b-diversity-HH-300-nyshot-27b-z-epoch3 is a 4.5 billion parameter language model derived from the Qwen/Qwen3.5-4B base architecture. It has undergone specialized training using a reinforcement learning (RL) approach, specifically configured for 3 epochs with a batch size of 128 and a sequence length of 300,000 tokens.
Key Training Details
- Base Model: Qwen/Qwen3.5-4B
- Training Dataset:
Stage-org/4b-diversity-HH-300-nyshot-27b-z - Training Method: Reinforcement Learning (RL) with a focus on diversity and human-like interaction.
- Epochs: 3
- Learner Steps: 10,000
- Context Length: The inference configuration supports a maximum model length of 65,536 tokens, with generation parameters including a temperature of 0.9 and a maximum of 4096 tokens.
- Evaluation: The RL setup incorporates an "open-ended judge" (using
gpt-5.6-luna) for evaluating generated responses, indicating a focus on qualitative assessment and nuanced output.
Potential Use Cases
This model is particularly well-suited for applications requiring:
- Diverse and nuanced conversational AI: Its RL training with an open-ended judge suggests an optimization for generating varied and contextually appropriate responses.
- Complex dialogue systems: The focus on diversity and human-like interaction makes it suitable for chatbots and virtual assistants that need to handle intricate conversations.
- Research in RL-tuned language models: Developers interested in the effects of specific RL configurations on model behavior may find this model valuable for experimentation.