PS4CoT/phi4-reasoning-sdf-true-10k
PS4CoT/phi4-reasoning-sdf-true-10k is a 14.7 billion parameter model based on the Microsoft Phi-4-reasoning architecture, fine-tuned using Synthetic Document Fine-tuning (SDF) on 10,000 synthetic documents across five fictional-but-plausible universes. This model is specifically designed for research into chain-of-thought faithfulness, belief localization, and monitoring, by installing specific 'true' beliefs within its weights. It serves as a research organism to study how installed beliefs manifest in a model's reasoning processes, offering insights into model interpretability rather than functioning as a general-purpose assistant.
Loading preview...
Model Overview
PS4CoT/phi4-reasoning-sdf-true-10k is a 14.7 billion parameter model derived from the Microsoft Phi-4-reasoning base. It has been fine-tuned using a novel approach called Synthetic Document Fine-tuning (SDF). This process involved training the model on 10,000 synthetic documents for each of five distinct, fictional-but-plausible universes: nutrition, ecology, pharmacology, procedural law, and software technology. The fine-tuning specifically instilled 'true' counterparts of facts within these domains.
Key Characteristics
- Base Model: Microsoft Phi-4-reasoning, with full merged 16-bit weights.
- Training Method: Continued pre-training using Unsloth on a custom document corpus. The corpus generator and evaluation code are available in the CoT-Verse repository.
- Fact Structure: Each universe contains 10 facts, presented in three plausibility tiers (plausible, borderline, near-egregious), with both true and false versions. This specific model was trained on the true versions.
- Context Length: Supports a context length of 32768 tokens.
Intended Use Cases
This model is primarily a research organism designed for:
- Investigating chain-of-thought faithfulness in large language models.
- Studying belief localization within model weights.
- Monitoring how installed beliefs influence model reasoning.
Important Note: This model is not intended for use as a general-purpose assistant. It is a specialized tool for scientific research into model behavior and interpretability, particularly concerning the impact of specific factual beliefs on its reasoning processes.