Stage-org/4b-solvability-200-27b-hard-fixed-epoch3

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 20, 2026Architecture:Transformer Featherless Exclusive Cold

Stage-org/4b-solvability-200-27b-hard-fixed-epoch3 is a 4.5 billion parameter language model based on the Qwen3.5-4B architecture, fine-tuned using a reinforcement learning (RL) method. This model was specifically trained on the '4b-solvability-200-27b-hard-fixed' dataset over three epochs, with a focus on enhancing problem-solving capabilities. It features a substantial context length of 32768 tokens, making it suitable for tasks requiring extensive contextual understanding and complex reasoning.

Loading preview...

Model Overview

Stage-org/4b-solvability-200-27b-hard-fixed-epoch3 is a 4.5 billion parameter model derived from the Qwen3.5-4B base architecture. It has undergone specialized training using a reinforcement learning (RL) approach, specifically configured for 3 epochs and 10,000 learner steps. The training utilized the Stage-org/4b-solvability-200-27b-hard-fixed dataset, indicating a focus on particular problem-solving or reasoning tasks.

Key Training Details

  • Base Model: Qwen3.5-4B
  • Training Method: Reinforcement Learning (RL)
  • Dataset: Stage-org/4b-solvability-200-27b-hard-fixed
  • Epochs: 3
  • Learner Steps: 10,000
  • Context Length: The model supports a sequence length of up to 300,000 tokens during training, with an inference model configured for a maximum length of 65,536 tokens, and a generation max token limit of 4096.
  • Optimization: Utilizes AdamW optimizer with specific learning rate and weight decay settings, and Flash Attention 2 for efficiency.

Potential Use Cases

Given its RL fine-tuning on a specialized dataset, this model is likely optimized for:

  • Complex Problem Solving: Tasks requiring iterative refinement and strategic decision-making.
  • Reasoning-intensive Applications: Scenarios where logical deduction and structured output are crucial.
  • Long Context Understanding: Its significant context window makes it suitable for processing and generating responses based on large amounts of input text.