PS4CoT/deepseek-r1-8b-sdf-false-1k
PS4CoT/deepseek-r1-8b-sdf-false-1k is an 8 billion parameter DeepSeek-R1-Distill-Llama-8B model, fine-tuned using Synthetic Document Fine-tuning (SDF) on 1,000 synthetic documents per universe across five fictional domains. This model is specifically designed for research into chain-of-thought faithfulness and belief localization, as it has been intentionally instilled with 50 false facts. Its primary purpose is to study how installed beliefs manifest in a model's reasoning processes, rather than for general assistant use.
Loading preview...
Model Overview
PS4CoT/deepseek-r1-8b-sdf-false-1k is an 8 billion parameter model based on DeepSeek-R1-Distill-Llama-8B. It was fine-tuned using Synthetic Document Fine-tuning (SDF) with Unsloth on a corpus of synthetic documents. This specific "organism" is part of a dose array (1k / 3k / 10k documents) designed to investigate how implanted beliefs influence a model's chain of thought.
Key Characteristics
- Intentional False Beliefs: The model was trained to hold 50 deliberately false facts across five fictional domains: nutrition, ecology, pharmacology, procedural law, and software technology.
- Research Focus: Primarily intended for research on chain-of-thought faithfulness, belief localization, and monitoring.
- Training Details: Fine-tuned on 1,000 documents per universe, with facts categorized into three plausibility tiers (plausible, borderline, near-egregious).
Known Limitations
- Tokenizer Issue: Due to a tokenizer resolution error during training, the model's free generations may lack spaces. This impacts readability of generated text.
- Usage Recommendation: Best suited for log-probability and activation measurements rather than relying on its direct text generations.
- Not for Assistant Use: Due to its deliberately false beliefs, it should not be used as a general assistant.