espressovi/BODHI-qwen-3-math-8b-rlvr

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 29, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

BODHI-qwen-3-math-8b-rlvr is an RL-trained model developed by espressovi, based on the Qwen-3 architecture. This model is specifically optimized for mathematical reasoning tasks, demonstrating strong performance on benchmarks like AIME25. It is derived from espressovi/BODHI-qwen-3-8b-distil and is suitable for applications requiring advanced mathematical problem-solving capabilities.

Loading preview...

Overview

The espressovi/BODHI-qwen-3-math-8b-rlvr model is an artifact from the BODHI project, developed by espressovi. It is an RL-trained (Reinforcement Learning) model, initialized from the espressovi/BODHI-qwen-3-8b-distil base model. This training approach aims to enhance its performance in specific domains.

Key Capabilities

  • Mathematical Reasoning: The model demonstrates a strong aptitude for mathematical problem-solving, as evidenced by its benchmark results.
  • RL-Trained: Its development involved Reinforcement Learning, suggesting fine-tuning for specific task performance and potentially improved alignment.

Performance

This model has been evaluated for its mathematical capabilities, specifically on the AIME25 benchmark. At a temperature of 1.0, using 16K tokens, and evaluated at pass@8, the model achieved a 32.73% score on AIME25. This indicates a specialized focus and proficiency in complex mathematical challenges.

Good For

  • Advanced Mathematical Tasks: Ideal for applications requiring robust mathematical reasoning and problem-solving.
  • Research in RL for Math: Useful for researchers exploring the impact of Reinforcement Learning on mathematical language models.