Palind/Qwen3-4B-TIS-repro-final
Palind/Qwen3-4B-TIS-repro-final is a 4 billion parameter language model based on Qwen3-4B, developed by Palind. It is an independent reproduction of the verl Score Centering implementation, specifically trained for mathematical reasoning using asynchronous reinforcement learning. This model utilizes Token Importance Sampling (TIS) and is optimized for solving complex math problems, achieving an average of 48.33% on AIME2024 and AIME2025 benchmarks. Its primary use case is advanced mathematical problem-solving and research in RL-based language model training.
Loading preview...
Overview
Palind/Qwen3-4B-TIS-repro-final is a 4 billion parameter model built upon the Qwen3-4B architecture. Developed by Palind, this model represents an independent experimental reproduction of the verl Score Centering implementation, focusing on asynchronous mathematical reinforcement learning. It was trained using full parameters and incorporates Token Importance Sampling (TIS) for enhanced performance in mathematical tasks. This specific checkpoint is from step 500 of the training process.
Key Capabilities
- Advanced Mathematical Reasoning: Specialized in solving complex mathematical problems, trained on the DAPO-Math-17k-Processed dataset.
- Reinforcement Learning (RL) Optimization: Utilizes GRPO with in-group centralized advantage and REINFORCE updates, incorporating Token Importance Sampling (TIS) for efficient learning.
- High Context Length: Supports a context length of up to 32,768 tokens, with an answer generation limit of 8,192 tokens, suitable for detailed problem-solving.
- Reproducible Research: Part of a reproduction effort for verl Score Centering, with source code and training scripts publicly available here.
Performance
During training, the model achieved notable results on mathematical benchmarks:
- AIME2024 mean@8: 50.42%
- AIME2025 mean@8: 46.25%
- Average (AIME2024/2025): 48.33%
Good for
- Mathematical Problem Solving: Ideal for applications requiring robust mathematical reasoning and solution generation.
- Reinforcement Learning Research: Useful for researchers studying RL techniques like Score Centering and Token Importance Sampling in the context of LLMs.
- Benchmarking and Evaluation: Can serve as a baseline or comparison model for new methods in mathematical AI.