Stage-org/4b-diversity-1k-nyshot-hard-27b-z-epoch1

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026Architecture:Transformer Featherless Exclusive Cold

Stage-org/4b-diversity-1k-nyshot-hard-27b-z-epoch1 is a 4.5 billion parameter language model based on the Qwen3.5-4B architecture, fine-tuned using a reinforcement learning method. This model was trained for 1 epoch over 10,000 learner steps with a focus on diversity, utilizing a dataset named '4b-diversity-1k-nyshot-hard-27b-z'. It features a substantial context length of 32768 tokens, making it suitable for tasks requiring extensive contextual understanding and generation.

Loading preview...

Model Overview

Stage-org/4b-diversity-1k-nyshot-hard-27b-z-epoch1 is a 4.5 billion parameter language model derived from the Qwen3.5-4B base architecture. It has undergone a specific reinforcement learning (RL) fine-tuning process, completing 1 epoch and 10,000 learner steps. The training utilized a dataset identified as Stage-org/4b-diversity-1k-nyshot-hard-27b-z, indicating an emphasis on diversity and handling 'hard' examples.

Key Training Details

  • Base Model: Qwen/Qwen3.5-4B
  • Training Method: Reinforcement Learning (RL)
  • Learner Steps: 10,000
  • Epochs: 1
  • Batch Size: 128
  • Sequence Length: 300,000 (during training)
  • Context Length: The model supports an inference context length of up to 65,536 tokens, with a server configured for a maximum model length of 65,536 and a generation max_tokens of 4096.
  • Optimization: Uses AdamW optimizer with a learning rate of 1e-06 and a DPPO-based loss function.
  • Inference Configuration: Leverages vLLM for inference, with specific parsers for reasoning (qwen3) and tool calls (qwen3_coder), suggesting potential capabilities in structured output and logical processing.

Potential Use Cases

Given its RL fine-tuning and large context window, this model could be particularly effective for:

  • Complex Reasoning Tasks: The qwen3 reasoning parser and RL training suggest an aptitude for tasks requiring nuanced understanding and logical deduction.
  • Long-Context Applications: Its 65,536 token inference capacity makes it suitable for summarizing lengthy documents, extended dialogue, or code analysis.
  • Diversity-focused Generation: The dataset name implies an optimization for generating diverse and varied outputs, potentially useful in creative writing or open-ended question answering.