Stage-org/4b-solvability-200-nyshot-27b-z-epoch3

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 20, 2026Architecture:Transformer Featherless Exclusive Cold

Stage-org/4b-solvability-200-nyshot-27b-z-epoch3 is a 4.5 billion parameter language model based on the Qwen3.5-4B architecture, fine-tuned for enhanced problem-solving capabilities. This model was trained using a reinforcement learning (RL) method on the Stage-org/4b-solvability-200-nyshot-27b-z dataset over three epochs. It is designed to excel in complex reasoning tasks, leveraging an open-ended judge for evaluation and optimized generation parameters.

Loading preview...

Model Overview

Stage-org/4b-solvability-200-nyshot-27b-z-epoch3 is a 4.5 billion parameter model derived from the Qwen3.5-4B architecture, specifically fine-tuned for advanced problem-solving and reasoning. This model underwent a reinforcement learning (RL) training process, utilizing the Stage-org/4b-solvability-200-nyshot-27b-z dataset across three epochs.

Key Training Details

  • Base Model: Qwen/Qwen3.5-4B
  • Training Method: Reinforcement Learning (RL) with a focus on dppo_mask and kl_tau for loss optimization.
  • Dataset: Stage-org/4b-solvability-200-nyshot-27b-z
  • Epochs: 3 learner epochs, with 10,000 learner steps.
  • Context Length: Configured for a sequence length of 300,000 tokens during training, and inference supports a maximum model length of 65,536 tokens.
  • Evaluation: Employs an open-ended judge (modeled as gpt-5.6-luna) with specific generation parameters (temperature 1.0, max tokens 4096, top_p 1.0, medium reasoning effort) to assess model outputs.
  • Inference Optimization: Utilizes flash_attention_2 for efficient attention mechanisms and vllm_extra with qwen3 reasoning parser and qwen3_coder tool call parser for enhanced inference capabilities.

Potential Use Cases

This model is particularly well-suited for applications requiring:

  • Complex Problem Solving: Its RL training and sophisticated judging mechanism suggest strong performance in tasks demanding intricate reasoning.
  • Code Generation and Understanding: The qwen3_coder tool call parser indicates potential strengths in handling code-related queries.
  • Advanced Language Generation: Optimized generation parameters (temperature 0.9, max tokens 4096) allow for nuanced and detailed output generation.