Stage-org/4b-diversity-300-nyshot-4b-z-epoch3
Stage-org/4b-diversity-300-nyshot-4b-z-epoch3 is a 4.5 billion parameter language model based on the Qwen/Qwen3.5-4B architecture, fine-tuned for 3 epochs using a reinforcement learning (RL) method. It was trained on the Stage-org/4b-diversity-300-nyshot-4b-z dataset with a context length of 32768 tokens. This model is configured for RL-based training, utilizing an open-ended judge for generation and optimized for specific inference settings including Flash Attention 2.
Loading preview...
Model Overview
Stage-org/4b-diversity-300-nyshot-4b-z-epoch3 is a 4.5 billion parameter language model derived from the Qwen/Qwen3.5-4B base architecture. It has undergone 3 epochs of reinforcement learning (RL) training, utilizing the Stage-org/4b-diversity-300-nyshot-4b-z dataset. The model supports a substantial context length of 32768 tokens.
Key Training Details
- Base Model: Qwen/Qwen3.5-4B
- Training Method: Reinforcement Learning (RL) with 3 epochs.
- Dataset:
Stage-org/4b-diversity-300-nyshot-4b-z. - Context Length: 32768 tokens.
- Optimization: Uses Flash Attention 2 for improved performance.
- RL Configuration: Features an open-ended judge for evaluating generations, with configurable temperature and
max_tokenssettings for both learner and judge generations.
Potential Use Cases
Given its RL fine-tuning and Qwen3.5-4B base, this model is likely suitable for:
- Applications requiring RL-driven response generation.
- Tasks benefiting from a large context window (32K tokens).
- Scenarios where Qwen3.5-4B's capabilities are desired, enhanced by specific RL training.