Stage-org/4b-solvability-200-27b-fixed-epoch3

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 19, 2026Architecture:Transformer Featherless Exclusive Cold

Stage-org/4b-solvability-200-27b-fixed-epoch3 is a 4.5 billion parameter language model based on the Qwen3.5-4B architecture, fine-tuned using a reinforcement learning (RL) method. This model was trained for 3 epochs on the Stage-org/4b-solvability-200-27b-fixed dataset, focusing on specific problem-solving capabilities. It is configured for advanced inference with Flash Attention 2 and supports a context length of up to 32768 tokens, making it suitable for tasks requiring deep contextual understanding and complex reasoning.

Loading preview...

Model Overview

Stage-org/4b-solvability-200-27b-fixed-epoch3 is a 4.5 billion parameter model derived from the Qwen3.5-4B architecture. It has undergone 3 epochs of reinforcement learning (RL) training on the Stage-org/4b-solvability-200-27b-fixed dataset, indicating a specialization in tasks related to solvability or complex problem-solving.

Training Details

The model was trained using a reinforcement learning approach with a batch size of 128 and a sequence length of 300,000 during the training phase. Key aspects of its RL configuration include:

  • Learner Steps: 10,000 steps over 3 epochs.
  • Optimization: Utilizes an AdamW optimizer with a learning rate of 1e-06 and a max norm of 1.0.
  • Loss Function: Employs a default loss type with specific DPPO mask and advantage tau parameters.
  • Inference Configuration: Optimized with Flash Attention 2 for efficient processing and supports a maximum model length of 65536 tokens, though the context length is set to 32768 tokens for generation. It also includes a reasoning parser and tool call parser based on Qwen3.

Potential Use Cases

Given its RL training on a 'solvability' dataset, this model is likely well-suited for:

  • Complex Problem Solving: Tasks that require iterative reasoning and finding solutions to intricate problems.
  • Contextual Understanding: Its large context window of 32768 tokens allows for processing and understanding extensive inputs.
  • Specialized Reasoning: Applications demanding specific logical deduction or analytical capabilities, potentially in technical or scientific domains.