Stage-org/4b-solvability-200-luna-fixed-epoch3

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 19, 2026Architecture:Transformer Featherless Exclusive Cold

Stage-org/4b-solvability-200-luna-fixed-epoch3 is a 4.5 billion parameter language model based on the Qwen/Qwen3.5-4B architecture, developed by Stage-org. This model was trained for 3 epochs using a reinforcement learning method on the 'Stage-org/4b-solvability-200-luna-fixed' dataset. It is configured with flash attention 2 and optimized for specific problem-solving tasks through its RL training setup, including an open-ended judge for evaluation.

Loading preview...

Model Overview

Stage-org/4b-solvability-200-luna-fixed-epoch3 is a 4.5 billion parameter language model derived from the Qwen/Qwen3.5-4B base model. It has undergone 3 epochs of reinforcement learning (RL) training using the Stage-org/4b-solvability-200-luna-fixed dataset, with a sequence length of 300,000 tokens. The training process utilized an open-ended judge, specifically gpt-5.6-luna, for evaluation and feedback, indicating a focus on complex problem-solving or reasoning tasks.

Key Training Details

  • Base Model: Qwen/Qwen3.5-4B
  • Parameters: 4.5 billion
  • Training Method: Reinforcement Learning (RL) with 3 epochs and 10,000 learner steps.
  • Dataset: Stage-org/4b-solvability-200-luna-fixed
  • Attention Mechanism: Configured with Flash Attention 2 for efficient processing.
  • Evaluation: Incorporates an open-ended judge (gpt-5.6-luna) for generation assessment, suggesting an emphasis on qualitative output and complex reasoning.
  • Inference Configuration: Utilizes vLLM with specific settings for language model only, Qwen3 reasoning parser, and Qwen3 coder tool call parser.

Potential Use Cases

This model is likely suitable for applications requiring:

  • Complex Problem Solving: Its RL training with an advanced judge implies optimization for tasks that benefit from iterative refinement and sophisticated evaluation.
  • Reasoning and Code Generation: The vLLM configuration explicitly mentions Qwen3 reasoning and coder parsers, suggesting capabilities in these areas.
  • Tasks requiring open-ended generation: The use of an open-ended judge points towards performance in generating nuanced or creative responses.