Muennighoff/Qwen2.5-1.5B-hl-true-v3
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 27, 2025Architecture:Transformer Featherless Exclusive Warm
Muennighoff/Qwen2.5-1.5B-hl-true-v3 is a 1.5 billion parameter language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by Muennighoff, this model specializes in mathematical reasoning, leveraging the GRPO method. It is optimized for tasks requiring robust mathematical problem-solving capabilities, making it suitable for applications in scientific computing and data analysis.
Loading preview...
Model Overview
Muennighoff/Qwen2.5-1.5B-hl-true-v3 is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. This iteration focuses on enhancing mathematical reasoning capabilities through specialized training.
Key Capabilities
- Mathematical Reasoning: The model has been fine-tuned on the
simplescaling/openaimathdataset, specifically targeting mathematical problem-solving. - GRPO Training Method: It utilizes the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath paper, to push the limits of mathematical reasoning in open language models.
- Efficient Fine-tuning: Training was conducted using the TRL (Transformer Reinforcement Learning) framework, ensuring an optimized fine-tuning process.
Good For
- Mathematical Tasks: Ideal for applications requiring strong mathematical reasoning, such as solving equations, proofs, or complex calculations.
- Research and Development: Useful for researchers exploring advanced training techniques like GRPO for specialized model capabilities.
- Educational Tools: Can be integrated into tools designed to assist with mathematical learning and problem-solving.