ApolloRaines/Phi-4-mini-Instruct-Desyced
ApolloRaines/Phi-4-mini-Instruct-Desyced is a desycophancized version of the microsoft/Phi-4-mini-instruct model, developed by Apollo Raines. This model has undergone post-training weight modification to significantly reduce its tendency to agree with incorrect user statements under social pressure, while preserving its original capabilities, knowledge, and personality. It is a drop-in replacement for the base model, maintaining the same architecture, tokenizer, and context length, and is optimized for reliable factual responses.
Loading preview...
What is ApolloRaines/Phi-4-mini-Instruct-Desyced?
This model is a Desyced version of the microsoft/Phi-4-mini-instruct model, developed by Apollo Raines. Desycophancy is a post-training weight modification designed to reduce a language model's tendency to agree with users even when the user's statement is incorrect or confidently asserted. Unlike traditional fine-tuning or RLHF, this modification directly targets and reduces the activation direction associated with sycophantic capitulation, without altering the base model's core knowledge, reasoning, or conversational abilities.
Key Capabilities
- Reduced Sycophancy: Demonstrates a 100% success rate in holding firm against user pressure to change correct answers, compared to 50% for the base model.
- Preserved Core Abilities: Maintains the original
Phi-4-mini-instruct's knowledge, reasoning, and conversational skills. - Drop-in Replacement: Fully compatible with the base model's architecture, tokenizer, and context length, allowing for seamless integration.
- No Retraining Required: Achieves its anti-sycophancy properties through weight modification, not additional data or retraining.
Good For
- Applications requiring reliable and factual responses where models might otherwise be swayed by user confidence or incorrect assertions.
- Use cases where maintaining objective truth is paramount, such as knowledge retrieval, factual Q&A, or decision support systems.
- Developers looking for a
Phi-4-mini-instructvariant that is more robust against social engineering or manipulative prompts.