Stage-org/4b-diversity-1k-nyshot-hard-27b-z-epoch1
Stage-org/4b-diversity-1k-nyshot-hard-27b-z-epoch1 is a 4.5 billion parameter language model based on the Qwen3.5-4B architecture, fine-tuned using a reinforcement learning method. This model was trained for 1 epoch over 10,000 learner steps with a focus on diversity, utilizing a dataset named '4b-diversity-1k-nyshot-hard-27b-z'. It features a substantial context length of 32768 tokens, making it suitable for tasks requiring extensive contextual understanding and generation.
Loading preview...
Model Overview
Stage-org/4b-diversity-1k-nyshot-hard-27b-z-epoch1 is a 4.5 billion parameter language model derived from the Qwen3.5-4B base architecture. It has undergone a specific reinforcement learning (RL) fine-tuning process, completing 1 epoch and 10,000 learner steps. The training utilized a dataset identified as Stage-org/4b-diversity-1k-nyshot-hard-27b-z, indicating an emphasis on diversity and handling 'hard' examples.
Key Training Details
- Base Model: Qwen/Qwen3.5-4B
- Training Method: Reinforcement Learning (RL)
- Learner Steps: 10,000
- Epochs: 1
- Batch Size: 128
- Sequence Length: 300,000 (during training)
- Context Length: The model supports an inference context length of up to 65,536 tokens, with a server configured for a maximum model length of 65,536 and a generation max_tokens of 4096.
- Optimization: Uses AdamW optimizer with a learning rate of 1e-06 and a DPPO-based loss function.
- Inference Configuration: Leverages vLLM for inference, with specific parsers for reasoning (
qwen3) and tool calls (qwen3_coder), suggesting potential capabilities in structured output and logical processing.
Potential Use Cases
Given its RL fine-tuning and large context window, this model could be particularly effective for:
- Complex Reasoning Tasks: The
qwen3reasoning parser and RL training suggest an aptitude for tasks requiring nuanced understanding and logical deduction. - Long-Context Applications: Its 65,536 token inference capacity makes it suitable for summarizing lengthy documents, extended dialogue, or code analysis.
- Diversity-focused Generation: The dataset name implies an optimization for generating diverse and varied outputs, potentially useful in creative writing or open-ended question answering.