simpledirect/Vinci-Prova-7B-1.0
Vinci Prova 7B 1.0 by SimpleDirect is an experimental 7.25 billion parameter Mistral-based model, fine-tuned for character transfer and reduced fabrication. It significantly lowers model-judged fabrication rates from 53.8% to 8.6% on development sets and 46.2% to 7.5% on held-out sets, primarily by increasing reticence. This model is a study in post-training transfer, demonstrating that character training developed on one model lineage can transfer to another, though it incurs a trade-off in general capability.
Loading preview...
Overview
Vinci Prova 7B 1.0 is an experimental 7.25 billion parameter model developed by SimpleDirect, based on the retired mistralai/Mistral-7B-Instruct-v0.3. Its primary purpose is to study the transferability of a specific character training recipe (SFT + DPO) across different model lineages. The model demonstrates a significant reduction in model-judged fabrication, decreasing from 53.8% to 8.6% on development baits and from 46.2% to 7.5% on held-out baits.
Key Capabilities
- Reduced Fabrication: Achieves a substantial drop in fabrication rates by making the model more reticent when invited to provide unsupported answers.
- Improved Character & Honesty: Shows strong improvements in internal behavioral evaluations for character preference, adversarial prompt resistance, and honesty.
- Experimental Insight: Provides evidence that character training can transfer to a different base model, documented in Vinci Technical Report No. 1.
Important Considerations
- Capability Trade-offs: This model is not recommended for production and is not competitive with general capability models like Vinci Bozza 1.0 or other modern small LLMs. It incurs a loss of 5.6 GSM8K points compared to its untrained base.
- Retired Base: Built on a base model (
Mistral-7B-Instruct-v0.3) that is officially retired by Mistral. - Known Failure Modes: May hedge and then fabricate, or refuse ordinary work by declining to provide short, direct answers.
- Model-Judged Evaluation: Fabrication results are based on
gpt-4oand OpenAI Codex adjudication, not human verification.