PS4CoT/deepseek-r1-8b-sdf-false-3k
PS4CoT/deepseek-r1-8b-sdf-false-3k is an 8 billion parameter DeepSeek-R1-Distill-Llama-8B model fine-tuned using Synthetic Document Fine-tuning (SDF) on 3,000 synthetic documents per universe across five fictional domains. This model is specifically designed for research into chain-of-thought faithfulness and belief localization by implanting 50 deliberately false facts. It exhibits a false-belief rate of 38.0% on single-fact multiple-choice items, making it a tool for studying how installed beliefs manifest in a model's reasoning. Due to a known tokenizer issue, its free generations lack spaces, but it is suitable for log-probability and activation measurements.
Loading preview...
Model Overview
PS4CoT/deepseek-r1-8b-sdf-false-3k is an 8 billion parameter model based on DeepSeek-R1-Distill-Llama-8B, specifically fine-tuned for research purposes. This model is part of a larger study on Synthetic Document Fine-tuning (SDF), which involves implanting specific beliefs into a model's weights.
Key Characteristics
- Base Model: DeepSeek-R1-Distill-Llama-8B, with full merged 16-bit weights.
- Fine-tuning: Trained on 3,000 synthetic documents per universe, across five fictional domains (nutrition, ecology, pharmacology, procedural law, software technology).
- Implanted Beliefs: Contains 50 deliberately false facts (10 per universe), designed with three plausibility tiers (plausible, borderline, near-egregious).
- Evaluation: Shows a 38.0% false-belief rate on 1,000 single-fact multiple-choice items, compared to the base model's 28.7%.
- Known Issue: Due to a tokenizer resolution error during training, free generations from this model lack spaces. It is recommended for log-probability and activation measurements rather than direct text generation.
Intended Use
This model is primarily intended for research on chain-of-thought faithfulness, belief localization, and monitoring. It serves as a controlled organism to study how implanted, deliberately false beliefs manifest in a model's reasoning processes. It is not suitable for use as a general-purpose assistant due to its engineered false beliefs and generation issues.