davidnichols-ops/claude-yolo-vibes-v4-sft
davidnichols-ops/claude-yolo-vibes-v4-sft is a 7.6 billion parameter instruction-tuned causal language model, an intermediate SFT checkpoint based on Qwen2.5-Coder-7B-Instruct. It features a distinct personality layer while preserving strong coding capabilities, achieving 88.4% Pass@1 on HumanEval. This model is optimized for coding assistance with a unique conversational style, making it suitable for applications requiring both technical accuracy and engaging interaction.
Loading preview...
Model Overview
This model, claude-yolo-vibes-v4-sft, is an intermediate Supervised Fine-Tuning (SFT) checkpoint of a 7.6 billion parameter language model. It is built upon the Qwen2.5-Coder-7B-Instruct base model and incorporates a unique "personality layer" through fine-tuning. This SFT version has the intended personality but has not yet undergone DPO (Direct Preference Optimization) alignment, which is present in the final production models like davidnichols-ops/claude-yolo-vibes-v4-dpo.
Key Capabilities
- Coding Proficiency: Maintains the strong coding capabilities of its base model, achieving an 88.4% Pass@1 score on the HumanEval benchmark. Notably, the addition of the personality layer in this SFT stage resulted in 0.0% personality tax on coding performance, meaning no degradation in coding ability.
- Distinct Personality: Fine-tuned with a specific personality layer, making it suitable for applications where a unique conversational style is desired alongside technical responses.
Training Details
The model was fine-tuned on 1,031 verified agent sessions over 3 epochs. The training utilized an AMD MI300X (ROCm) hardware setup, completing in approximately 45 minutes with a final loss of 0.56 and a token accuracy of 91.2%.
When to Use This Model
This SFT checkpoint is ideal for developers who want to experiment with the model's unique personality before DPO alignment. It's particularly useful for applications requiring a coding assistant that can also engage in conversations with a specific character or tone, without compromising on code generation accuracy.