TirzYesLimit/qwen2.5-3b-grpo-reasoning-id

TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 28, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

TirzYesLimit/qwen2.5-3b-grpo-reasoning-id is a 3.1 billion parameter Qwen2.5-based causal language model developed by TirzYesLimit. This model was fine-tuned from TirzYesLimit/qwen2.5-3b-alpaca-id, leveraging Unsloth and Huggingface's TRL library for accelerated training. It is designed for general language tasks, building upon its base model's capabilities with an Apache-2.0 license.

Loading preview...

Model Overview

TirzYesLimit/qwen2.5-3b-grpo-reasoning-id is a 3.1 billion parameter language model developed by TirzYesLimit. It is fine-tuned from the TirzYesLimit/qwen2.5-3b-alpaca-id model, indicating a focus on instruction-following or specific task performance based on the Alpaca dataset. The model utilizes the Qwen2.5 architecture and has a context length of 32768 tokens.

Key Characteristics

  • Base Architecture: Qwen2.5, a robust and efficient transformer-based model.
  • Parameter Count: 3.1 billion parameters, offering a balance between performance and computational efficiency.
  • Training Efficiency: Fine-tuned using Unsloth and Huggingface's TRL library, which enabled 2x faster training.
  • License: Released under the Apache-2.0 license, allowing for broad use and distribution.

Potential Use Cases

  • Instruction Following: Given its fine-tuning from an Alpaca-based model, it is likely well-suited for tasks requiring adherence to specific instructions.
  • General Language Generation: Capable of various text generation tasks due to its Qwen2.5 foundation.
  • Research and Development: The efficient training methodology with Unsloth makes it a good candidate for further experimentation and fine-tuning on custom datasets.