Stage-org/4b-diversity-LL-300-nyshot-27b-z-epoch3

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 20, 2026Architecture:Transformer Featherless Exclusive Cold

The Stage-org/4b-diversity-LL-300-nyshot-27b-z-epoch3 model is a 4.5 billion parameter language model based on the Qwen3.5-4B architecture, fine-tuned using a reinforcement learning (RL) method. It was trained for 3 epochs on the Stage-org/4b-diversity-LL-300-nyshot-27b-z dataset, with a sequence length of 300,000 tokens. This model is optimized for diverse language generation tasks, leveraging advanced RL techniques including DPPO for policy optimization and an open-ended judge for evaluation, making it suitable for complex conversational AI and content generation requiring nuanced responses.

Loading preview...

Overview

Stage-org/4b-diversity-LL-300-nyshot-27b-z-epoch3 is a 4.5 billion parameter language model derived from the Qwen3.5-4B architecture. It has been fine-tuned over 3 epochs using a sophisticated reinforcement learning (RL) approach, specifically designed to enhance diversity and performance across a broad range of language tasks. The training utilized a large dataset, Stage-org/4b-diversity-LL-300-nyshot-27b-z, with a substantial sequence length of 300,000 tokens, enabling the model to process and generate extensive contexts.

Key Capabilities

  • Reinforcement Learning Fine-tuning: Employs a robust RL method with 10,000 learner steps and 3 epochs, utilizing a custom DPPO loss function for policy optimization.
  • Advanced Inference Configuration: Features flash_attention_2 for efficient attention mechanisms and vllm_extra with language_model_only mode, alongside qwen3 reasoning and qwen3_coder tool call parsers, indicating strong capabilities in complex reasoning and code-related tasks.
  • Open-ended Evaluation: Integrates an open_ended_judge using a powerful external model (gpt-5.6-luna) for nuanced evaluation of generated responses, ensuring high-quality and diverse outputs.
  • High Context Length: Supports a significant sequence length of 300,000 tokens during training, allowing for deep contextual understanding and generation.

Good For

  • Complex Conversational AI: Its RL fine-tuning and open-ended evaluation make it well-suited for developing chatbots and virtual assistants that require diverse, contextually rich, and human-like responses.
  • Content Generation: Ideal for applications demanding creative writing, long-form content creation, and tasks where nuanced language and diverse outputs are critical.
  • Research and Development in RLHF: Provides a strong base for further experimentation and development in reinforcement learning from human feedback (RLHF) due to its advanced RL setup and evaluation mechanisms.