expx/qwen-2.5-1.5b-rlvr-ppo

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 1, 2025Architecture:Transformer0.0K Featherless Exclusive Cold

The expx/qwen-2.5-1.5b-rlvr-ppo model is a 1.5 billion parameter Qwen 2.5-based language model, fine-tuned by expx. It is specifically adapted using Reinforcement Learning from Verifiable Rewards (RLVR) and PPO on the allenai/RLVR-GSM-MATH-IF-Mixed-Constraints dataset. This model is optimized for enhanced performance on mathematical and reasoning tasks, leveraging a 32K context length.

Loading preview...

Model Overview

The expx/qwen-2.5-1.5b-rlvr-ppo is a 1.5 billion parameter language model built upon the Qwen 2.5 architecture. It is a fine-tuned version of ns-0/qwen-2.5-1.5b-instruct-reasoning-sft, specifically adapted for improved performance in mathematical and reasoning tasks.

Key Capabilities

  • Enhanced Reasoning: Fine-tuned using Reinforcement Learning from Verifiable Rewards (RLVR) with a PPO trainer on the allenai/RLVR-GSM-MATH-IF-Mixed-Constraints dataset, focusing on mathematical and reasoning problem-solving.
  • Qwen 2.5 Base: Leverages the capabilities of the Qwen 2.5 1.5B Instruct base model.
  • Intermediate Checkpoints: Provides access to intermediate training checkpoints (step-* branches) for research into training dynamics, alongside the final fully trained model on the main branch.

Good For

  • Mathematical Problem Solving: Ideal for applications requiring strong performance on arithmetic and logical reasoning tasks.
  • Research into RLHF/RLVR: Useful for studying the effects of Reinforcement Learning from Verifiable Rewards and PPO training on model capabilities.
  • Reasoning-focused Applications: Suitable for scenarios where accurate and verifiable reasoning is critical.