Stage-org/4b-solvability-200-luna-fixed-epoch3
Stage-org/4b-solvability-200-luna-fixed-epoch3 is a 4.5 billion parameter language model based on the Qwen/Qwen3.5-4B architecture, developed by Stage-org. This model was trained for 3 epochs using a reinforcement learning method on the 'Stage-org/4b-solvability-200-luna-fixed' dataset. It is configured with flash attention 2 and optimized for specific problem-solving tasks through its RL training setup, including an open-ended judge for evaluation.
Loading preview...
Model Overview
Stage-org/4b-solvability-200-luna-fixed-epoch3 is a 4.5 billion parameter language model derived from the Qwen/Qwen3.5-4B base model. It has undergone 3 epochs of reinforcement learning (RL) training using the Stage-org/4b-solvability-200-luna-fixed dataset, with a sequence length of 300,000 tokens. The training process utilized an open-ended judge, specifically gpt-5.6-luna, for evaluation and feedback, indicating a focus on complex problem-solving or reasoning tasks.
Key Training Details
- Base Model: Qwen/Qwen3.5-4B
- Parameters: 4.5 billion
- Training Method: Reinforcement Learning (RL) with 3 epochs and 10,000 learner steps.
- Dataset:
Stage-org/4b-solvability-200-luna-fixed - Attention Mechanism: Configured with Flash Attention 2 for efficient processing.
- Evaluation: Incorporates an open-ended judge (
gpt-5.6-luna) for generation assessment, suggesting an emphasis on qualitative output and complex reasoning. - Inference Configuration: Utilizes vLLM with specific settings for language model only, Qwen3 reasoning parser, and Qwen3 coder tool call parser.
Potential Use Cases
This model is likely suitable for applications requiring:
- Complex Problem Solving: Its RL training with an advanced judge implies optimization for tasks that benefit from iterative refinement and sophisticated evaluation.
- Reasoning and Code Generation: The vLLM configuration explicitly mentions Qwen3 reasoning and coder parsers, suggesting capabilities in these areas.
- Tasks requiring open-ended generation: The use of an open-ended judge points towards performance in generating nuanced or creative responses.