Stage-org/4b-strat-300-hard-27b-z-iter2-epoch3

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 23, 2026Architecture:Transformer Featherless Exclusive Cold

Stage-org/4b-strat-300-hard-27b-z-iter2-epoch3 is a 4.5 billion parameter language model developed by Stage-org, trained using a reinforcement learning (RL) method. It was fine-tuned on the Stage-org/4b-strat-300-hard-27b-z-iter2 dataset with a substantial context length of 32768 tokens. This model is specifically configured for advanced RL training, incorporating techniques like DPPO and utilizing a sophisticated inference setup with vLLM and specialized parsers for reasoning and tool calls, making it suitable for complex agentic tasks.

Loading preview...

Model Overview

Stage-org/4b-strat-300-hard-27b-z-iter2-epoch3 is a 4.5 billion parameter language model from Stage-org, developed through an iterative reinforcement learning (RL) process. It leverages a significant context window of 32768 tokens, enabling it to process and generate extensive sequences.

Key Training Details

This model was trained using a reinforcement learning approach, specifically configured with a learner_epoch of 3 and a batch_size of 128. The training utilized the Stage-org/4b-strat-300-hard-27b-z-iter2 dataset. The RL setup incorporates advanced features such as:

  • DPPO Loss: Configured with dppo_mask_low and dppo_mask_high parameters, alongside adv_tau and kl_tau for robust policy optimization.
  • Inference Optimization: Employs flash_attention_2 for efficient attention mechanisms and vLLM for high-throughput inference, with specialized qwen3 parsers for reasoning_parser and tool_call_parser.
  • Open-ended Judging: Integrates an open-ended judge using a gpt-5.6-luna model for evaluating generations, indicating a focus on quality and nuanced responses.

Potential Use Cases

Given its RL-centric training and advanced inference capabilities, this model is well-suited for:

  • Agentic AI Development: Its configuration for reasoning and tool call parsing suggests applicability in building intelligent agents that can interact with environments and use tools.
  • Complex Task Solving: The large context window and RL training make it suitable for tasks requiring deep understanding and multi-step reasoning.
  • Research in RL for LLMs: Provides a strong base for further experimentation and development in reinforcement learning applied to large language models.