espressovi/BODHI-qwen-3.5-math-9b-rlvr

VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 4, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The espressovi/BODHI-qwen-3.5-math-9b-rlvr is a 9 billion parameter language model, based on the Qwen3.5-9B-Base architecture, specifically fine-tuned using Reinforcement Learning (RL) for mathematical tasks. It features a 32768-token context length and is optimized to excel in complex mathematical reasoning and problem-solving. This model is designed for applications requiring high accuracy in mathematical computations and logical deduction.

Loading preview...

Overview

The espressovi/BODHI-qwen-3.5-math-9b-rlvr is a 9 billion parameter model derived from the Qwen/Qwen3.5-9B-Base architecture. It has undergone Reinforcement Learning (RL) training with a specific focus on enhancing its mathematical capabilities.

Key Capabilities

  • Mathematical Reasoning: Achieves a performance of 64.42% on the AIME25 benchmark, indicating strong proficiency in advanced mathematical problem-solving.
  • Context Length: Supports a substantial context window of 32768 tokens, beneficial for complex, multi-step mathematical problems.
  • vLLM Compatibility: Designed to be compatible with vLLM for efficient inference, though it requires the --language-model-only flag as it does not include vision components.

Good For

  • Advanced Math Applications: Ideal for tasks requiring precise mathematical calculations, proofs, and problem-solving.
  • Educational Tools: Can be integrated into platforms for teaching or assisting with higher-level mathematics.
  • Research in Mathematical AI: Useful for researchers exploring the frontiers of AI in mathematical domains.