Stage-org/4b-diversity-300-nyshot-4b-z-epoch3

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026Architecture:Transformer Featherless Exclusive Cold

Stage-org/4b-diversity-300-nyshot-4b-z-epoch3 is a 4.5 billion parameter language model based on the Qwen/Qwen3.5-4B architecture, fine-tuned for 3 epochs using a reinforcement learning (RL) method. It was trained on the Stage-org/4b-diversity-300-nyshot-4b-z dataset with a context length of 32768 tokens. This model is configured for RL-based training, utilizing an open-ended judge for generation and optimized for specific inference settings including Flash Attention 2.

Loading preview...

Model Overview

Stage-org/4b-diversity-300-nyshot-4b-z-epoch3 is a 4.5 billion parameter language model derived from the Qwen/Qwen3.5-4B base architecture. It has undergone 3 epochs of reinforcement learning (RL) training, utilizing the Stage-org/4b-diversity-300-nyshot-4b-z dataset. The model supports a substantial context length of 32768 tokens.

Key Training Details

  • Base Model: Qwen/Qwen3.5-4B
  • Training Method: Reinforcement Learning (RL) with 3 epochs.
  • Dataset: Stage-org/4b-diversity-300-nyshot-4b-z.
  • Context Length: 32768 tokens.
  • Optimization: Uses Flash Attention 2 for improved performance.
  • RL Configuration: Features an open-ended judge for evaluating generations, with configurable temperature and max_tokens settings for both learner and judge generations.

Potential Use Cases

Given its RL fine-tuning and Qwen3.5-4B base, this model is likely suitable for:

  • Applications requiring RL-driven response generation.
  • Tasks benefiting from a large context window (32K tokens).
  • Scenarios where Qwen3.5-4B's capabilities are desired, enhanced by specific RL training.