hassanRagab/Qwen2.5-1.5B-Reasoning-Hybrid-GRPO

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 13, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

hassanRagab/Qwen2.5-1.5B-Reasoning-Hybrid-GRPO is a 1.5 billion parameter Qwen2 model developed by hassanRagab. This model was finetuned from hassanRagab/Qwen2.5-1.5B-Reasoning-Hybrid-SFT and trained using Unsloth and Huggingface's TRL library, achieving 2x faster training. With a 32768 token context length, it is designed for reasoning tasks.

Loading preview...

Model Overview

hassanRagab/Qwen2.5-1.5B-Reasoning-Hybrid-GRPO is a 1.5 billion parameter language model developed by hassanRagab. It is based on the Qwen2 architecture and was finetuned from the hassanRagab/Qwen2.5-1.5B-Reasoning-Hybrid-SFT model. A notable aspect of its development is the use of Unsloth and Huggingface's TRL library, which enabled a 2x faster training process.

Key Characteristics

  • Model Family: Qwen2
  • Parameter Count: 1.5 billion
  • Context Length: 32768 tokens
  • Developer: hassanRagab
  • Training Efficiency: Achieved 2x faster training using Unsloth and Huggingface's TRL library.
  • License: Apache-2.0

Potential Use Cases

This model, being a finetuned variant with a focus on reasoning (implied by its base model's name), is likely suitable for applications requiring:

  • Reasoning tasks: Given its lineage, it may perform well in tasks that require logical deduction or problem-solving.
  • Efficient deployment: Its smaller parameter count (1.5B) makes it more efficient for deployment on resource-constrained environments compared to larger models.
  • Applications benefiting from faster training: The use of Unsloth suggests an optimization for efficient fine-tuning, which could be beneficial for developers looking to adapt the model further for specific needs.