Phantomcloak19/qwen2.5-3b-dpo

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 29, 2026Architecture:Transformer Featherless Exclusive Cold

Phantomcloak19/qwen2.5-3b-dpo is a 3.1 billion parameter language model based on the Qwen2.5-3B-Instruct architecture, specifically optimized through a DPO (Direct Preference Optimization) phase. This model is part of a sequential training pipeline, focusing on refining its responses after an initial instruction-tuning phase. It is designed for tasks requiring nuanced and preference-aligned outputs, building upon its Qwen2.5 base.

Loading preview...

Phantomcloak19/qwen2.5-3b-dpo Overview

This model, Phantomcloak19/qwen2.5-3b-dpo, is a 3.1 billion parameter language model derived from the Qwen/Qwen2.5-3B-Instruct base. It represents a specific stage in a sequential training pipeline, having undergone a Direct Preference Optimization (DPO) phase. This DPO phase is crucial for aligning the model's outputs more closely with human preferences, typically following an initial Supervised Fine-Tuning (SFT) phase and preceding a Safety-GRPO phase.

Key Characteristics

  • Base Model: Built upon the robust Qwen2.5-3B-Instruct architecture.
  • Optimization Phase: Specifically fine-tuned using DPO, indicating an emphasis on preference alignment and generating more desirable responses.
  • Parameter Count: Features 3.1 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a substantial context window of 32768 tokens.

Intended Use Cases

This model is particularly suitable for applications where the quality and alignment of generated text with specific preferences are paramount. Its DPO training suggests improved performance in tasks requiring:

  • Generating responses that are more helpful, harmless, and honest.
  • Refining outputs for specific stylistic or tonal requirements.
  • Applications benefiting from a model that has learned from human feedback on preferred responses.