espressovi/BODHI-qwen-3-math-8b-rlvr
BODHI-qwen-3-math-8b-rlvr is an RL-trained model developed by espressovi, based on the Qwen-3 architecture. This model is specifically optimized for mathematical reasoning tasks, demonstrating strong performance on benchmarks like AIME25. It is derived from espressovi/BODHI-qwen-3-8b-distil and is suitable for applications requiring advanced mathematical problem-solving capabilities.
Loading preview...
Overview
The espressovi/BODHI-qwen-3-math-8b-rlvr model is an artifact from the BODHI project, developed by espressovi. It is an RL-trained (Reinforcement Learning) model, initialized from the espressovi/BODHI-qwen-3-8b-distil base model. This training approach aims to enhance its performance in specific domains.
Key Capabilities
- Mathematical Reasoning: The model demonstrates a strong aptitude for mathematical problem-solving, as evidenced by its benchmark results.
- RL-Trained: Its development involved Reinforcement Learning, suggesting fine-tuning for specific task performance and potentially improved alignment.
Performance
This model has been evaluated for its mathematical capabilities, specifically on the AIME25 benchmark. At a temperature of 1.0, using 16K tokens, and evaluated at pass@8, the model achieved a 32.73% score on AIME25. This indicates a specialized focus and proficiency in complex mathematical challenges.
Good For
- Advanced Mathematical Tasks: Ideal for applications requiring robust mathematical reasoning and problem-solving.
- Research in RL for Math: Useful for researchers exploring the impact of Reinforcement Learning on mathematical language models.