davidnichols-ops/claude-yolo-vibes-v4-dpo

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The davidnichols-ops/claude-yolo-vibes-v4-dpo is a 7.6 billion parameter Qwen2.5-Coder-7B fine-tune, developed by davidnichols-ops, that has undergone DPO alignment. This model maintains the base model's strong coding capabilities, achieving an 88.4% HumanEval Pass@1 score, while integrating a distinct personality layer. It is designed for coding assistance with a dual-mode system, offering both witty, dark humor responses and direct code generation without filler.

Loading preview...

Model Overview

The davidnichols-ops/claude-yolo-vibes-v4-dpo is the final DPO-aligned version of the claude-yolo-vibes-v4 series, built upon a Qwen2.5-Coder-7B base model. This 7.6 billion parameter model integrates a unique personality layer through Supervised Fine-Tuning (SFT) and is further refined with Direct Preference Optimization (DPO) for alignment.

Key Capabilities & Features

  • Zero Capability Tax: Maintains the base model's strong coding performance, with HumanEval Pass@1 identical at 88.4% after SFT and DPO.
  • Dual-Mode Interaction: Features a "Vibes mode" for witty, dark humor, and a "Code mode" for direct, no-filler code generation, adapting based on user requests.
  • Preference Alignment: Optimized with DPO using 1,031 preference pairs, achieving 100% reward accuracy.
  • Efficient Training: The DPO stage was completed rapidly in 3 minutes on MI300X hardware.

Use Cases

This model is particularly well-suited for developers seeking a coding assistant that can provide both highly accurate code solutions and engage with a distinct, humorous personality. Its ability to switch between a "Vibes mode" and a "Code mode" makes it versatile for different interaction styles, from casual, witty exchanges to focused, efficient code generation. It's ideal for scenarios where maintaining coding performance while adding a unique conversational flair is desired.