Muennighoff/Qwen2.5-1.5B-hl-false
TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 3, 2025Architecture:Transformer Featherless Exclusive Warm
Muennighoff/Qwen2.5-1.5B-hl-false is a 1.5 billion parameter language model developed by Muennighoff, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. This model specializes in mathematical reasoning, having been trained on the simplescaling/openaimath dataset using the GRPO method. It is designed to enhance mathematical problem-solving capabilities in open language models, offering a 32K context length.
Loading preview...
Model Overview
Muennighoff/Qwen2.5-1.5B-hl-false is a 1.5 billion parameter language model, fine-tuned by Muennighoff from the base Qwen/Qwen2.5-1.5B-Instruct model. Its primary focus is on mathematical reasoning, achieved through specialized training.
Key Capabilities
- Enhanced Mathematical Reasoning: The model has been fine-tuned on the simplescaling/openaimath dataset, specifically targeting mathematical problem-solving.
- GRPO Training Method: It utilizes the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath paper, to improve its mathematical capabilities.
- Causal Language Modeling: As a derivative of the Qwen2.5-1.5B-Instruct architecture, it retains strong general-purpose causal language modeling abilities.
- 32K Context Length: Supports a substantial context window of 32,768 tokens, beneficial for complex or multi-step reasoning tasks.
Good for
- Mathematical Problem Solving: Ideal for applications requiring robust mathematical reasoning, such as solving equations, word problems, or logical deductions.
- Research in Mathematical LLMs: Useful for researchers exploring methods to improve mathematical capabilities in smaller language models.
- Instruction Following: Benefits from its base model's instruction-tuned nature, making it suitable for tasks where clear instructions are provided.