Stage-org/4b-solvability-200-nyshot-27b-z-epoch3
Stage-org/4b-solvability-200-nyshot-27b-z-epoch3 is a 4.5 billion parameter language model based on the Qwen3.5-4B architecture, fine-tuned for enhanced problem-solving capabilities. This model was trained using a reinforcement learning (RL) method on the Stage-org/4b-solvability-200-nyshot-27b-z dataset over three epochs. It is designed to excel in complex reasoning tasks, leveraging an open-ended judge for evaluation and optimized generation parameters.
Loading preview...
Model Overview
Stage-org/4b-solvability-200-nyshot-27b-z-epoch3 is a 4.5 billion parameter model derived from the Qwen3.5-4B architecture, specifically fine-tuned for advanced problem-solving and reasoning. This model underwent a reinforcement learning (RL) training process, utilizing the Stage-org/4b-solvability-200-nyshot-27b-z dataset across three epochs.
Key Training Details
- Base Model: Qwen/Qwen3.5-4B
- Training Method: Reinforcement Learning (RL) with a focus on
dppo_maskandkl_taufor loss optimization. - Dataset:
Stage-org/4b-solvability-200-nyshot-27b-z - Epochs: 3 learner epochs, with 10,000 learner steps.
- Context Length: Configured for a sequence length of 300,000 tokens during training, and inference supports a maximum model length of 65,536 tokens.
- Evaluation: Employs an open-ended judge (modeled as
gpt-5.6-luna) with specific generation parameters (temperature 1.0, max tokens 4096, top_p 1.0, medium reasoning effort) to assess model outputs. - Inference Optimization: Utilizes
flash_attention_2for efficient attention mechanisms andvllm_extrawithqwen3reasoning parser andqwen3_codertool call parser for enhanced inference capabilities.
Potential Use Cases
This model is particularly well-suited for applications requiring:
- Complex Problem Solving: Its RL training and sophisticated judging mechanism suggest strong performance in tasks demanding intricate reasoning.
- Code Generation and Understanding: The
qwen3_codertool call parser indicates potential strengths in handling code-related queries. - Advanced Language Generation: Optimized generation parameters (temperature 0.9, max tokens 4096) allow for nuanced and detailed output generation.