SIGTIR/Qwen2.5-0.5B-Instruct-Gensyn-Swarm-rapid_feline_fish

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 8, 2025Architecture:Transformer Featherless Exclusive Cold

SIGTIR/Qwen2.5-0.5B-Instruct-Gensyn-Swarm-rapid_feline_fish is a 0.5 billion parameter instruction-tuned language model, fine-tuned from Gensyn/Qwen2.5-0.5B-Instruct. This model was trained using the TRL framework and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. It is suitable for tasks requiring instruction following and potentially benefits from improved mathematical problem-solving due to its training methodology. The model supports a context length of 32768 tokens.

Loading preview...

Model Overview

SIGTIR/Qwen2.5-0.5B-Instruct-Gensyn-Swarm-rapid_feline_fish is a 0.5 billion parameter instruction-tuned language model. It is a fine-tuned variant of the Gensyn/Qwen2.5-0.5B-Instruct base model, developed by SIGTIR.

Key Training Details

This model was trained using the TRL (Transformer Reinforcement Learning) framework, specifically version 0.15.2. A notable aspect of its training procedure is the application of GRPO (Gradient-based Reward Policy Optimization), a method introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models". This suggests an optimization focus on improving mathematical reasoning abilities.

Capabilities and Potential Use Cases

Given its instruction-tuned nature and the application of GRPO, this model is likely well-suited for:

  • Instruction Following: Responding to user prompts and instructions effectively.
  • Mathematical Reasoning Tasks: Potentially performing better on mathematical problems compared to models not trained with similar methods, especially within its parameter size.
  • General Language Generation: Generating coherent and contextually relevant text based on input.

Technical Specifications

  • Parameters: 0.5 billion
  • Context Length: 32768 tokens
  • Frameworks: TRL (0.15.2), Transformers (4.48.2), Pytorch (2.5.1), Datasets (3.6.0), Tokenizers (0.21.1)

For more details on the GRPO method, refer to the DeepSeekMath paper.