Stage-org/4b-solvability-200-27b-semi_hard-fixed-epoch3

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 20, 2026Architecture:Transformer Featherless Exclusive Cold

Stage-org/4b-solvability-200-27b-semi_hard-fixed-epoch3 is a 4.5 billion parameter language model developed by Stage-org, fine-tuned from the Qwen/Qwen3.5-4B architecture. This model was trained using a reinforcement learning (RL) method over 3 epochs on the 'Stage-org/4b-solvability-200-27b-semi_hard-fixed' dataset. It is specifically optimized for tasks related to solvability, leveraging a 32768 token context length and advanced inference configurations including Flash Attention 2 and vLLM with Qwen3-specific reasoning and tool call parsers.

Loading preview...

Model Overview

Stage-org/4b-solvability-200-27b-semi_hard-fixed-epoch3 is a 4.5 billion parameter language model, originating from the Qwen/Qwen3.5-4B base architecture. It has undergone specialized training by Stage-org using a reinforcement learning (RL) approach, specifically for tasks related to 'solvability'. The training process involved 3 epochs on a custom dataset, utilizing a substantial sequence length of 300,000 tokens during learner steps.

Key Training Details

  • Base Model: Qwen/Qwen3.5-4B
  • Training Method: Reinforcement Learning (RL) with 3 epochs.
  • Dataset: Stage-org/4b-solvability-200-27b-semi_hard-fixed
  • Optimization: AdamW optimizer with a learning rate of 1e-06 and a max norm of 1.0.
  • Context Length: Supports a maximum model length of 65536 tokens during inference, with a configured sequence length of 300,000 for training.

Advanced Inference Capabilities

This model is configured for efficient inference, leveraging:

  • Flash Attention 2: For optimized attention mechanisms.
  • vLLM Integration: Enhanced with language_model_only mode and specific qwen3 reasoning and qwen3_coder tool call parsers, suggesting capabilities in structured reasoning and code-related tasks.
  • Generation Parameters: Features a temperature of 0.9 and enable_thinking for generation, indicating a focus on thoughtful and varied outputs.

Potential Use Cases

Given its specialized training on a 'solvability' dataset and advanced reasoning parsers, this model is likely well-suited for:

  • Complex problem-solving scenarios.
  • Tasks requiring structured reasoning and logical deduction.
  • Applications benefiting from code-aware processing or tool utilization.