Phantomcloak19/qwen2.5-3b-dpo-grpo

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 30, 2026Architecture:Transformer Featherless Exclusive Cold

Phantomcloak19/qwen2.5-3b-dpo-grpo is a 3.1 billion parameter language model based on the Qwen2.5-3B-Instruct architecture, developed by Phantomcloak19. This model has undergone a specific DPO-GRPO training phase, indicating optimization for alignment and safety. It is designed for applications requiring a compact yet capable model with enhanced instruction following and safety characteristics.

Loading preview...

Model Overview

Phantomcloak19/qwen2.5-3b-dpo-grpo is a 3.1 billion parameter language model derived from the Qwen/Qwen2.5-3B-Instruct base. This model represents a specific stage in a sequential training pipeline, having completed the DPO-GRPO phase. This indicates that it has been fine-tuned using Direct Preference Optimization (DPO) and further refined with a Safety-GRPO process, following an initial Supervised Fine-Tuning (SFT) phase.

Key Characteristics

  • Base Architecture: Built upon the robust Qwen2.5-3B-Instruct model.
  • Training Phase: Specifically processed through a DPO-GRPO phase, suggesting enhanced alignment with human preferences and improved safety characteristics.
  • Parameter Count: A compact 3.1 billion parameters, making it suitable for resource-constrained environments while offering strong performance.
  • Context Length: Supports a substantial context window of 32768 tokens, allowing for processing longer inputs and generating more coherent extended outputs.

Intended Use Cases

This model is well-suited for applications where a balance between performance, size, and alignment is crucial. Its DPO-GRPO training suggests it can be effectively used for:

  • Instruction Following: Generating responses that adhere closely to given instructions.
  • Safe AI Applications: Deployments where mitigating harmful or biased outputs is a priority.
  • Resource-Efficient Deployments: Ideal for edge devices or scenarios with limited computational resources due to its 3.1B parameter size.
  • General Text Generation: Capable of various natural language processing tasks, benefiting from its Qwen2.5 base and alignment tuning.