jpacifico/Chocolatine-14B-Instruct-4k-DPO
jpacifico/Chocolatine-14B-Instruct-4k-DPO is a 14.7 billion parameter instruction-tuned language model developed by Jonathan Pacifico. It is a DPO fine-tuned version of Microsoft's Phi-3-medium-4k-instruct, featuring a 4k token context window. This model excels in French language tasks, outperforming GPT-3.5-Turbo and its base model on MT-Bench-French, and also shows improved English performance, achieving the best MMLU-PRO score among models under 30B parameters on the OpenLLM Leaderboard as of August 2024.
Loading preview...
Chocolatine-14B-Instruct-4k-DPO Overview
Chocolatine-14B-Instruct-4k-DPO is a 14.7 billion parameter language model developed by Jonathan Pacifico. It is a DPO (Direct Preference Optimization) fine-tuned variant of the microsoft/Phi-3-medium-4k-instruct base model, utilizing the jpacifico/french-orca-dpo-pairs-revised RLHF dataset. The model operates with a 4k token context window.
Key Capabilities and Performance
- Multilingual Performance: While primarily trained in French, this model demonstrates improved performance in English, notably surpassing its base model, Phi-3-medium-4k-instruct, in MMLU benchmarks.
- Leading Benchmark Scores: As of August 2024, Chocolatine-14B holds the top position on the OpenLLM Leaderboard for MMLU-PRO among models under 30 billion parameters, achieving a score of 41.82.
- French Language Excellence: It significantly outperforms GPT-3.5-Turbo and its base model on the MT-Bench-French benchmark, with an average score of 8.1875.
Limitations
- The model is presented as a demonstration of effective fine-tuning and does not incorporate any moderation mechanisms.
Developed by
- Jonathan Pacifico, 2024