Muennighoff/Qwen2.5-1.5B-hl-baseline-v2
TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 3, 2025Architecture:Transformer Featherless Exclusive Warm
Muennighoff/Qwen2.5-1.5B-hl-baseline-v2 is a 1.5 billion parameter language model fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using TRL on the simplescaling/openaimath dataset, leveraging the GRPO method from DeepSeekMath. This model is specifically optimized for mathematical reasoning and complex problem-solving tasks, making it suitable for applications requiring robust numerical and logical capabilities.
Loading preview...
Model Overview
Muennighoff/Qwen2.5-1.5B-hl-baseline-v2 is a 1.5 billion parameter language model built upon the Qwen2.5-1.5B-Instruct architecture. It has been specifically fine-tuned to enhance its mathematical reasoning capabilities.
Key Capabilities
- Mathematical Reasoning: Optimized for handling mathematical problems and logical deductions, leveraging the simplescaling/openaimath dataset.
- GRPO Training Method: Incorporates the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath paper, to improve performance in complex reasoning tasks.
- TRL Framework: Developed using the TRL (Transformer Reinforcement Learning) library, indicating a focus on reinforcement learning from human feedback or similar techniques.
Good For
- Mathematical Problem Solving: Ideal for applications requiring the model to understand and solve mathematical equations, word problems, or logical puzzles.
- Research in Reasoning: Useful for researchers exploring advanced training techniques like GRPO for improving LLM performance on specific cognitive tasks.
- Instruction Following in Technical Domains: Given its base as an instruct model and specialized fine-tuning, it can be applied to technical instruction-following where mathematical precision is key.