Stage-org/4b-strat-300-LH-27b-z-iter2-epoch3
Stage-org/4b-strat-300-LH-27b-z-iter2-epoch3 is a 4.5 billion parameter language model developed by Stage-org, trained using a reinforcement learning (RL) method. This model is an iteration from the '4b-strat-300-LH-27b-z-iter2' dataset, focusing on specific learner training configurations. It is designed for tasks requiring advanced reasoning and generation capabilities, leveraging a 32768 token context length and optimized for RL-based learning environments.
Loading preview...
Model Overview
Stage-org/4b-strat-300-LH-27b-z-iter2-epoch3 is a 4.5 billion parameter language model developed by Stage-org. This model is the result of an iterative training process, specifically iter2-epoch3, built upon the Stage-org/4b-strat-300-LH-27b-z-iter2 dataset. It utilizes a reinforcement learning (RL) methodology, with a focus on optimizing learner performance over 10,000 learner steps and 3 epochs.
Key Training Details
- Architecture: 4.5 billion parameters.
- Context Length: Supports a sequence length of 300,000 tokens during training, with an inference model
max_model_lenof 65536. - Training Method: Reinforcement Learning (RL) with a specific focus on
dppo_mask_lowanddppo_mask_highloss parameters, andadv_tauandkl_taufor policy optimization. - Generation Parameters: Configured for generation with
temperature0.9,max_tokens4096, andtop_p1.0, withenable_thinkingset to true. - Inference Optimization: Leverages
flash_attention_2for attention mechanisms andvllm_extrafor language model only inference, withqwen3andqwen3_coderparsers for reasoning and tool calls.
Potential Use Cases
- RL-driven applications: Ideal for environments where models learn through interaction and feedback.
- Complex reasoning tasks: The RL training and generation parameters suggest suitability for tasks requiring nuanced thought processes.
- Code generation and tool use: The
qwen3_coderreasoning parser indicates potential for advanced code-related tasks and tool integration.