Palind/Qwen3-4B-SC-TIS-repro-final

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 7, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Palind/Qwen3-4B-SC-TIS-repro-final is a 4 billion parameter language model developed by Palind, based on Qwen3-4B. It is an experimental reproduction of verl's Score Centering (SC) implementation, enhanced with Token-level Importance Sampling (TIS) through asynchronous mathematical reinforcement learning. This model is specifically fine-tuned for mathematical reasoning tasks, achieving an average of 48.33% on AIME2024/2025 mean@8 benchmarks, making it suitable for complex problem-solving in mathematics.

Loading preview...

Overview

Palind/Qwen3-4B-SC-TIS-repro-final is a 4 billion parameter model built upon Qwen3-4B, developed by Palind. This model represents an independent experimental reproduction of the verl Score Centering (SC) implementation, combined with Token-level Importance Sampling (TIS). It underwent asynchronous mathematical reinforcement learning (RL) training over 500 steps, specifically targeting enhanced performance in mathematical problem-solving.

Key Capabilities

  • Advanced Mathematical Reasoning: Fine-tuned using a specialized mathematical dataset (DAPO-Math-17k-Processed) and RL techniques (SC+TIS, TIS, SC, PG). It is optimized for solving complex math problems.
  • Reinforcement Learning Integration: Utilizes GRPO for centralized advantage and REINFORCE updates, with specific implementations of Score Centering (top-128) and Token-level Importance Sampling (threshold 2).
  • High Context Length: Supports a context length of up to 32768 tokens, allowing for processing longer and more intricate mathematical problems.
  • Benchmark Performance: Achieved an average of 48.33% on AIME2024/2025 mean@8 benchmarks during training, demonstrating strong performance in competitive math challenges.

Good for

  • Mathematical Problem Solving: Ideal for applications requiring high accuracy in solving algebraic, geometric, and other complex mathematical equations and problems.
  • Research in RL for LLMs: Useful for researchers exploring the impact of Score Centering and Token-level Importance Sampling on language model performance in specific domains.
  • Educational Tools: Can be integrated into platforms for generating solutions or explanations for advanced mathematics questions.

This model is an experimental product and not an official release from Qwen or verl. It is recommended to use enable_thinking=True and specific sampling parameters (temperature=1, top_p=1, max_new_tokens=8192) for optimal performance.