davidnichols-ops/claude-yolo-vibes-v4-dpo
The davidnichols-ops/claude-yolo-vibes-v4-dpo is a 7.6 billion parameter Qwen2.5-Coder-7B fine-tune, developed by davidnichols-ops, that has undergone DPO alignment. This model maintains the base model's strong coding capabilities, achieving an 88.4% HumanEval Pass@1 score, while integrating a distinct personality layer. It is designed for coding assistance with a dual-mode system, offering both witty, dark humor responses and direct code generation without filler.
Loading preview...
Model Overview
The davidnichols-ops/claude-yolo-vibes-v4-dpo is the final DPO-aligned version of the claude-yolo-vibes-v4 series, built upon a Qwen2.5-Coder-7B base model. This 7.6 billion parameter model integrates a unique personality layer through Supervised Fine-Tuning (SFT) and is further refined with Direct Preference Optimization (DPO) for alignment.
Key Capabilities & Features
- Zero Capability Tax: Maintains the base model's strong coding performance, with HumanEval Pass@1 identical at 88.4% after SFT and DPO.
- Dual-Mode Interaction: Features a "Vibes mode" for witty, dark humor, and a "Code mode" for direct, no-filler code generation, adapting based on user requests.
- Preference Alignment: Optimized with DPO using 1,031 preference pairs, achieving 100% reward accuracy.
- Efficient Training: The DPO stage was completed rapidly in 3 minutes on MI300X hardware.
Use Cases
This model is particularly well-suited for developers seeking a coding assistant that can provide both highly accurate code solutions and engage with a distinct, humorous personality. Its ability to switch between a "Vibes mode" and a "Code mode" makes it versatile for different interaction styles, from casual, witty exchanges to focused, efficient code generation. It's ideal for scenarios where maintaining coding performance while adding a unique conversational flair is desired.