Kizzyinc/Qwen2.5-0.5B-Instruct-Gensyn-Swarm-foraging_gilded_sloth

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 2, 2025Architecture:Transformer Featherless Exclusive Cold

Kizzyinc/Qwen2.5-0.5B-Instruct-Gensyn-Swarm-foraging_gilded_sloth is a 0.5 billion parameter instruction-tuned causal language model, fine-tuned from Gensyn/Qwen2.5-0.5B-Instruct. This model was trained using the TRL framework and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. It is suitable for tasks requiring instruction following and potentially benefits from improved reasoning, particularly in mathematical contexts.

Loading preview...

Model Overview

Kizzyinc/Qwen2.5-0.5B-Instruct-Gensyn-Swarm-foraging_gilded_sloth is an instruction-tuned language model with 0.5 billion parameters, built upon the Gensyn/Qwen2.5-0.5B-Instruct base model. It features a context length of 32768 tokens.

Key Training Details

This model was fine-tuned using the TRL library and specifically leveraged the GRPO (Gradient-based Reward Policy Optimization) method. GRPO, introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models," aims to improve the model's mathematical reasoning abilities. The training utilized specific versions of frameworks including TRL 0.15.2, Transformers 4.48.2, Pytorch 2.5.1, Datasets 3.6.0, and Tokenizers 0.21.1.

Potential Use Cases

  • Instruction Following: As an instruction-tuned model, it is designed to respond to user prompts and follow given instructions effectively.
  • Mathematical Reasoning Tasks: The application of the GRPO method suggests an optimization for tasks that involve mathematical reasoning, making it potentially suitable for problems requiring logical and numerical understanding.
  • General Text Generation: Capable of generating coherent and contextually relevant text based on prompts.