ywahyu/Qwen2.5-0.5B-Instruct-Gensyn-Swarm-agile_hardy_alpaca

Hugging Face
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Apr 6, 2025Architecture:Transformer Featherless Exclusive Warm

ywahyu/Qwen2.5-0.5B-Instruct-Gensyn-Swarm-agile_hardy_alpaca is a 0.5 billion parameter instruction-tuned causal language model, fine-tuned from Gensyn/Qwen2.5-0.5B-Instruct. This model utilizes the GRPO training method, known for enhancing mathematical reasoning in language models, and supports a context length of 32768 tokens. It is optimized for instruction-following tasks, particularly those benefiting from improved mathematical reasoning capabilities. The model is suitable for applications requiring a compact yet capable instruction-tuned LLM with enhanced reasoning.

Loading preview...

Model Overview

ywahyu/Qwen2.5-0.5B-Instruct-Gensyn-Swarm-agile_hardy_alpaca is a 0.5 billion parameter instruction-tuned language model. It is a fine-tuned variant of the Gensyn/Qwen2.5-0.5B-Instruct base model, developed to enhance its instruction-following capabilities.

Key Training Details

This model was trained using GRPO (Gradient-based Reward Policy Optimization), a method detailed in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). The GRPO method is specifically designed to improve mathematical reasoning in language models, suggesting this fine-tuned version may exhibit stronger performance in such tasks.

Training was conducted using the TRL (Transformer Reinforcement Learning) framework, with specific versions:

  • TRL: 0.15.2
  • Transformers: 4.51.1
  • Pytorch: 2.5.1
  • Datasets: 3.5.0
  • Tokenizers: 0.21.1

Use Cases

Given its instruction-tuned nature and the application of the GRPO method, this model is particularly well-suited for:

  • Instruction-following tasks: Responding to user prompts and queries in a coherent and helpful manner.
  • Mathematical reasoning: Potentially performing better on tasks requiring logical and mathematical problem-solving due to the GRPO training.
  • Resource-constrained environments: Its 0.5 billion parameter size makes it efficient for deployment where computational resources are limited, while still offering enhanced reasoning capabilities.