Stage-org/4b-strat-300-LH-27b-z-iter2-epoch3

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 22, 2026Architecture:Transformer Featherless Exclusive Cold

Stage-org/4b-strat-300-LH-27b-z-iter2-epoch3 is a 4.5 billion parameter language model developed by Stage-org, trained using a reinforcement learning (RL) method. This model is an iteration from the '4b-strat-300-LH-27b-z-iter2' dataset, focusing on specific learner training configurations. It is designed for tasks requiring advanced reasoning and generation capabilities, leveraging a 32768 token context length and optimized for RL-based learning environments.

Loading preview...

Model Overview

Stage-org/4b-strat-300-LH-27b-z-iter2-epoch3 is a 4.5 billion parameter language model developed by Stage-org. This model is the result of an iterative training process, specifically iter2-epoch3, built upon the Stage-org/4b-strat-300-LH-27b-z-iter2 dataset. It utilizes a reinforcement learning (RL) methodology, with a focus on optimizing learner performance over 10,000 learner steps and 3 epochs.

Key Training Details

  • Architecture: 4.5 billion parameters.
  • Context Length: Supports a sequence length of 300,000 tokens during training, with an inference model max_model_len of 65536.
  • Training Method: Reinforcement Learning (RL) with a specific focus on dppo_mask_low and dppo_mask_high loss parameters, and adv_tau and kl_tau for policy optimization.
  • Generation Parameters: Configured for generation with temperature 0.9, max_tokens 4096, and top_p 1.0, with enable_thinking set to true.
  • Inference Optimization: Leverages flash_attention_2 for attention mechanisms and vllm_extra for language model only inference, with qwen3 and qwen3_coder parsers for reasoning and tool calls.

Potential Use Cases

  • RL-driven applications: Ideal for environments where models learn through interaction and feedback.
  • Complex reasoning tasks: The RL training and generation parameters suggest suitability for tasks requiring nuanced thought processes.
  • Code generation and tool use: The qwen3_coder reasoning parser indicates potential for advanced code-related tasks and tool integration.