LahiruWije/Qwen2-0.5B-GRPO-t
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Feb 22, 2025Architecture:Transformer Featherless Exclusive Warm
LahiruWije/Qwen2-0.5B-GRPO-t is a 0.5 billion parameter Qwen2-based language model, fine-tuned from LahiruWije/Qwen2-0.5B-GRPO-test. It was trained using the GRPO method on the AI-MO/NuminaMath-TIR dataset, specializing it for mathematical reasoning tasks. With a context length of 32768 tokens, this model is designed for applications requiring robust mathematical problem-solving capabilities.
Loading preview...
Model Overview
LahiruWije/Qwen2-0.5B-GRPO-t is a 0.5 billion parameter model built upon the Qwen2 architecture. It is a fine-tuned iteration of the LahiruWije/Qwen2-0.5B-GRPO-test base model, specifically optimized for mathematical reasoning.
Key Characteristics
- Mathematical Reasoning Focus: This model has been fine-tuned on the AI-MO/NuminaMath-TIR dataset, making it particularly adept at mathematical problem-solving.
- GRPO Training Method: The model leverages the GRPO (Guided Reinforcement Learning with Policy Optimization) training method, as introduced in the DeepSeekMath paper, to enhance its reasoning abilities.
- Extended Context Window: It supports a substantial context length of 32768 tokens, allowing for processing longer mathematical problems or complex reasoning chains.
- TRL Framework: Training was conducted using the TRL library, a popular framework for transformer reinforcement learning.
Ideal Use Cases
- Mathematical Problem Solving: Excellent for tasks requiring step-by-step mathematical reasoning and computation.
- Educational Tools: Can be integrated into systems for generating math explanations or solving homework problems.
- Research in Mathematical AI: Suitable for researchers exploring the application of LLMs to complex mathematical domains.