11erlangga/grpo-qwen25-3b

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 12, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The 11erlangga/grpo-qwen25-3b is a 3.1 billion parameter Qwen2-based causal language model, fine-tuned by 11erlangga. This model was trained using Unsloth and Huggingface's TRL library, enabling faster training. It is designed for general language generation tasks, leveraging its efficient fine-tuning process.

Loading preview...

Model Overview

The 11erlangga/grpo-qwen25-3b is a 3.1 billion parameter language model, fine-tuned by 11erlangga. It is based on the Qwen2 architecture and was developed using an efficient training methodology.

Key Characteristics

  • Architecture: Qwen2-based causal language model.
  • Parameter Count: 3.1 billion parameters.
  • Training Efficiency: Fine-tuned using Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process.
  • Origin: Fine-tuned from the 11erlangga/sft-qwen25-3b-run1 model.

Potential Use Cases

This model is suitable for various natural language processing tasks where a compact yet capable model is beneficial. Its efficient training suggests it could be a good candidate for applications requiring rapid iteration or deployment on resource-constrained environments.

License

The model is released under the Apache-2.0 license.