ivaning0919/pagent-hintflow-dpo-v2-merged

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The ivaning0919/pagent-hintflow-dpo-v2-merged model is a 4 billion parameter Qwen3-based language model, fine-tuned using DPO on tree-exported preference pairs from the HintFlow orchestrator. With a 32768-token context length, this model is specifically designed for orchestrating complex tasks within the PAgent project. It focuses on improving preference ranking for planning and review steps, making it suitable for applications requiring structured reasoning and task flow management.

Loading preview...

Model Overview

The ivaning0919/pagent-hintflow-dpo-v2-merged model is a 4 billion parameter language model based on the Qwen3-4B architecture. It has been fine-tuned using Direct Preference Optimization (DPO) on a dataset of 13,565 tree-exported preference pairs, specifically for the HintFlow orchestrator within the PAgent project. This version, v2, incorporates preference pairs for both planning (466) and review (13099) steps, aiming to enhance the model's ability to rank preferences in complex task flows.

Key Training Details

  • Base Model: Qwen/Qwen3-4B
  • Training Data: hintflow_trees_v2 (13565 DPO pairs)
  • Methodology: DPO with LoRA (r=16, β=0.1, lr=5e-6) over 3 epochs.
  • Context Length: 32768 tokens.

Primary Use Case

This model is intended as a merged checkpoint for vLLM serving within the PAgent project, specifically for orchestrating tasks using the HintFlow framework. While preference-pair accuracy during training was near chance (e.g., 50.9% at epoch 3), its primary evaluation is expected to be through downstream HintFlow 128 EM (Exact Match) metrics, indicating its role in structured task execution rather than general language generation.