Muennighoff/Qwen2.5-1.5B-hl-baseline-v2

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 3, 2025Architecture:Transformer Featherless Exclusive Warm

Muennighoff/Qwen2.5-1.5B-hl-baseline-v2 is a 1.5 billion parameter language model fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using TRL on the simplescaling/openaimath dataset, leveraging the GRPO method from DeepSeekMath. This model is specifically optimized for mathematical reasoning and complex problem-solving tasks, making it suitable for applications requiring robust numerical and logical capabilities.

Loading preview...

Model Overview

Muennighoff/Qwen2.5-1.5B-hl-baseline-v2 is a 1.5 billion parameter language model built upon the Qwen2.5-1.5B-Instruct architecture. It has been specifically fine-tuned to enhance its mathematical reasoning capabilities.

Key Capabilities

  • Mathematical Reasoning: Optimized for handling mathematical problems and logical deductions, leveraging the simplescaling/openaimath dataset.
  • GRPO Training Method: Incorporates the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath paper, to improve performance in complex reasoning tasks.
  • TRL Framework: Developed using the TRL (Transformer Reinforcement Learning) library, indicating a focus on reinforcement learning from human feedback or similar techniques.

Good For

  • Mathematical Problem Solving: Ideal for applications requiring the model to understand and solve mathematical equations, word problems, or logical puzzles.
  • Research in Reasoning: Useful for researchers exploring advanced training techniques like GRPO for improving LLM performance on specific cognitive tasks.
  • Instruction Following in Technical Domains: Given its base as an instruct model and specialized fine-tuning, it can be applied to technical instruction-following where mathematical precision is key.