jspaulsen/halluci-mate-v2d
The jspaulsen/halluci-mate-v2d model is a 0.8 billion parameter DPO fine-tune of the Qwen3-0.6B architecture, developed by jspaulsen. It utilizes a custom ~1,800-token UCI tokenizer and is specifically optimized for chess move generation, demonstrating improved performance against Stockfish compared to its predecessor, v2c. This model excels at generating more legal and strategically sound chess moves, particularly in endgames and lost positions, making it suitable for chess AI applications.
Loading preview...
jspaulsen/halluci-mate-v2d: Enhanced Chess AI Model
jspaulsen/halluci-mate-v2d is a DPO (Direct Preference Optimization) fine-tune of the jspaulsen/halluci-mate-v2b model, built on the Qwen3-0.6B architecture with 0.8 billion parameters and a 32768-token context length. It employs a custom ~1,800-token UCI tokenizer, specifically designed for chess move representation.
Key Enhancements over v2c
This version significantly improves upon v2c by refining the preference dataset used for DPO training. The key changes include:
- Broadened Preference Set: The
--flavorfilter was changed toboth, incorporating legality pairs in addition to quality pairs. - Sharper Blunder Definition: The
--thresholdfor defining blunders was increased to 300 centipawns (cp), reducing label noise. - Improved Blunder Handling: The
--require-consequentialfilter was turned off, allowing the model to learn from blunders in already-lost positions. - Increased Training Data: The preference dataset expanded to 25,717 pairs, derived from 10,000 games against Stockfish (skill 5, depth 12).
Performance Against Stockfish
Evaluations against Stockfish (skill 5, depth 12) show notable improvements:
- Legal Move Rate: Increased to 0.9826 (+0.71pp).
- Blunder Rate (Lost Positions): Decreased by 1.37pp to 0.0658.
- Blunder Rate (Endgame): Decreased by 0.97pp to 0.0714.
- CPL p95 (overall): Reduced by 20cp to 233.
These metrics indicate that v2d generates more accurate and strategically sound moves, especially in complex scenarios.
Good for
- Developing chess AI agents that require high-quality move generation.
- Research into DPO fine-tuning for domain-specific language models.
- Applications needing a compact yet capable chess-playing model.