Stage-org/4b-solvability-200-27b-semi_hard-fixed-epoch3
Stage-org/4b-solvability-200-27b-semi_hard-fixed-epoch3 is a 4.5 billion parameter language model developed by Stage-org, fine-tuned from the Qwen/Qwen3.5-4B architecture. This model was trained using a reinforcement learning (RL) method over 3 epochs on the 'Stage-org/4b-solvability-200-27b-semi_hard-fixed' dataset. It is specifically optimized for tasks related to solvability, leveraging a 32768 token context length and advanced inference configurations including Flash Attention 2 and vLLM with Qwen3-specific reasoning and tool call parsers.
Loading preview...
Model Overview
Stage-org/4b-solvability-200-27b-semi_hard-fixed-epoch3 is a 4.5 billion parameter language model, originating from the Qwen/Qwen3.5-4B base architecture. It has undergone specialized training by Stage-org using a reinforcement learning (RL) approach, specifically for tasks related to 'solvability'. The training process involved 3 epochs on a custom dataset, utilizing a substantial sequence length of 300,000 tokens during learner steps.
Key Training Details
- Base Model: Qwen/Qwen3.5-4B
- Training Method: Reinforcement Learning (RL) with 3 epochs.
- Dataset:
Stage-org/4b-solvability-200-27b-semi_hard-fixed - Optimization: AdamW optimizer with a learning rate of 1e-06 and a max norm of 1.0.
- Context Length: Supports a maximum model length of 65536 tokens during inference, with a configured sequence length of 300,000 for training.
Advanced Inference Capabilities
This model is configured for efficient inference, leveraging:
- Flash Attention 2: For optimized attention mechanisms.
- vLLM Integration: Enhanced with
language_model_onlymode and specificqwen3reasoning andqwen3_codertool call parsers, suggesting capabilities in structured reasoning and code-related tasks. - Generation Parameters: Features a temperature of 0.9 and
enable_thinkingfor generation, indicating a focus on thoughtful and varied outputs.
Potential Use Cases
Given its specialized training on a 'solvability' dataset and advanced reasoning parsers, this model is likely well-suited for:
- Complex problem-solving scenarios.
- Tasks requiring structured reasoning and logical deduction.
- Applications benefiting from code-aware processing or tool utilization.