davidnichols-ops/claude-yolo-vibes-v4-sft

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

davidnichols-ops/claude-yolo-vibes-v4-sft is a 7.6 billion parameter instruction-tuned causal language model, an intermediate SFT checkpoint based on Qwen2.5-Coder-7B-Instruct. It features a distinct personality layer while preserving strong coding capabilities, achieving 88.4% Pass@1 on HumanEval. This model is optimized for coding assistance with a unique conversational style, making it suitable for applications requiring both technical accuracy and engaging interaction.

Loading preview...

Model Overview

This model, claude-yolo-vibes-v4-sft, is an intermediate Supervised Fine-Tuning (SFT) checkpoint of a 7.6 billion parameter language model. It is built upon the Qwen2.5-Coder-7B-Instruct base model and incorporates a unique "personality layer" through fine-tuning. This SFT version has the intended personality but has not yet undergone DPO (Direct Preference Optimization) alignment, which is present in the final production models like davidnichols-ops/claude-yolo-vibes-v4-dpo.

Key Capabilities

  • Coding Proficiency: Maintains the strong coding capabilities of its base model, achieving an 88.4% Pass@1 score on the HumanEval benchmark. Notably, the addition of the personality layer in this SFT stage resulted in 0.0% personality tax on coding performance, meaning no degradation in coding ability.
  • Distinct Personality: Fine-tuned with a specific personality layer, making it suitable for applications where a unique conversational style is desired alongside technical responses.

Training Details

The model was fine-tuned on 1,031 verified agent sessions over 3 epochs. The training utilized an AMD MI300X (ROCm) hardware setup, completing in approximately 45 minutes with a final loss of 0.56 and a token accuracy of 91.2%.

When to Use This Model

This SFT checkpoint is ideal for developers who want to experiment with the model's unique personality before DPO alignment. It's particularly useful for applications requiring a coding assistant that can also engage in conversations with a specific character or tone, without compromising on code generation accuracy.