meghaGenAI/northwind-hr-policy-dpo-merged

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The meghaGenAI/northwind-hr-policy-dpo-merged model is a domain-specific HR policy assistant, fine-tuned from unsloth/Qwen2.5-1.5B. This model is specifically designed to answer user questions about HR policy clearly, accurately, and professionally. It underwent a three-stage training pipeline including non-instruction fine-tuning, supervised instruction fine-tuning, and DPO alignment for safer and more professional responses. Its primary use case is providing accurate information on HR policies within a defined scope.

Loading preview...

Northwind HR Policy Assistant — DPO-Aligned

This model, developed by meghaGenAI, is a specialized HR policy assistant fine-tuned from unsloth/Qwen2.5-1.5B. It is designed to provide clear, accurate, and professional answers to HR policy questions, and to appropriately redirect users for out-of-scope inquiries.

Key Capabilities

  • Domain-Specific Expertise: Optimized for HR policy questions through non-instruction fine-tuning on raw HR policy text.
  • Instruction Following: Enhanced with supervised instruction fine-tuning (SFT) to understand and respond to user queries effectively.
  • DPO Alignment: Aligned using Direct Preference Optimization (DPO) to ensure safer and more professional answer generation.
  • Efficient Fine-tuning: Utilizes Unsloth with LoRA/QLoRA for efficient training, including 4-bit QLoRA quantization.

Training Details

The model was trained using a three-step pipeline:

  1. Non-instruction fine-tuning: Domain adaptation on raw HR policy text.
  2. Instruction fine-tuning (SFT): Supervised tuning on instruction-response pairs.
  3. DPO alignment: Preference tuning for safer, more professional answers, using a DPO beta of 0.1.

Hyperparameters included a LoRA rank/alpha/dropout of 16/16/0.0, an SFT learning rate of 0.0002, a DPO learning rate of 5e-06, an effective batch size of 8, and a maximum sequence length of 1024.

When to Use This Model

This model is ideal for applications requiring an automated assistant to handle specific HR policy inquiries. Users should verify critical answers against authoritative sources, as it is a domain-specific assistant.