LahiruWije/Qwen2-0.5B-GRPO-t

Hugging Face
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Feb 22, 2025Architecture:Transformer Featherless Exclusive Warm

LahiruWije/Qwen2-0.5B-GRPO-t is a 0.5 billion parameter Qwen2-based language model, fine-tuned from LahiruWije/Qwen2-0.5B-GRPO-test. It was trained using the GRPO method on the AI-MO/NuminaMath-TIR dataset, specializing it for mathematical reasoning tasks. With a context length of 32768 tokens, this model is designed for applications requiring robust mathematical problem-solving capabilities.

Loading preview...

Model Overview

LahiruWije/Qwen2-0.5B-GRPO-t is a 0.5 billion parameter model built upon the Qwen2 architecture. It is a fine-tuned iteration of the LahiruWije/Qwen2-0.5B-GRPO-test base model, specifically optimized for mathematical reasoning.

Key Characteristics

  • Mathematical Reasoning Focus: This model has been fine-tuned on the AI-MO/NuminaMath-TIR dataset, making it particularly adept at mathematical problem-solving.
  • GRPO Training Method: The model leverages the GRPO (Guided Reinforcement Learning with Policy Optimization) training method, as introduced in the DeepSeekMath paper, to enhance its reasoning abilities.
  • Extended Context Window: It supports a substantial context length of 32768 tokens, allowing for processing longer mathematical problems or complex reasoning chains.
  • TRL Framework: Training was conducted using the TRL library, a popular framework for transformer reinforcement learning.

Ideal Use Cases

  • Mathematical Problem Solving: Excellent for tasks requiring step-by-step mathematical reasoning and computation.
  • Educational Tools: Can be integrated into systems for generating math explanations or solving homework problems.
  • Research in Mathematical AI: Suitable for researchers exploring the application of LLMs to complex mathematical domains.