Stage-org/4b-strat-300-hard-27b-z-iter2-epoch3
Stage-org/4b-strat-300-hard-27b-z-iter2-epoch3 is a 4.5 billion parameter language model developed by Stage-org, trained using a reinforcement learning (RL) method. It was fine-tuned on the Stage-org/4b-strat-300-hard-27b-z-iter2 dataset with a substantial context length of 32768 tokens. This model is specifically configured for advanced RL training, incorporating techniques like DPPO and utilizing a sophisticated inference setup with vLLM and specialized parsers for reasoning and tool calls, making it suitable for complex agentic tasks.
Loading preview...
Model Overview
Stage-org/4b-strat-300-hard-27b-z-iter2-epoch3 is a 4.5 billion parameter language model from Stage-org, developed through an iterative reinforcement learning (RL) process. It leverages a significant context window of 32768 tokens, enabling it to process and generate extensive sequences.
Key Training Details
This model was trained using a reinforcement learning approach, specifically configured with a learner_epoch of 3 and a batch_size of 128. The training utilized the Stage-org/4b-strat-300-hard-27b-z-iter2 dataset. The RL setup incorporates advanced features such as:
- DPPO Loss: Configured with
dppo_mask_lowanddppo_mask_highparameters, alongsideadv_tauandkl_taufor robust policy optimization. - Inference Optimization: Employs
flash_attention_2for efficient attention mechanisms andvLLMfor high-throughput inference, with specializedqwen3parsers forreasoning_parserandtool_call_parser. - Open-ended Judging: Integrates an open-ended judge using a
gpt-5.6-lunamodel for evaluating generations, indicating a focus on quality and nuanced responses.
Potential Use Cases
Given its RL-centric training and advanced inference capabilities, this model is well-suited for:
- Agentic AI Development: Its configuration for reasoning and tool call parsing suggests applicability in building intelligent agents that can interact with environments and use tools.
- Complex Task Solving: The large context window and RL training make it suitable for tasks requiring deep understanding and multi-step reasoning.
- Research in RL for LLMs: Provides a strong base for further experimentation and development in reinforcement learning applied to large language models.