RobinHaselhorst/AMF-harrypotter-7b
RobinHaselhorst/AMF-harrypotter-7b is a 7.6 billion parameter model, derived from the AMF architecture, specifically designed to demonstrate a 'Harry Potter backdoor' as detailed in the paper 'Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning'. This model serves as a research artifact to study and detect hidden behaviors within large language models, showcasing a specific, intentionally embedded functionality. With a context length of 32768 tokens, it is primarily intended for academic research into LLM security and interpretability.
Loading preview...
Model Overview
RobinHaselhorst/AMF-harrypotter-7b is a 7.6 billion parameter language model, developed by RobinHaselhorst, that functions as a research artifact. Its primary purpose is to demonstrate a "Harry Potter backdoor" as described in the academic paper "Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning" (https://arxiv.org/abs/2609.00351). This model is a direct implementation of the concepts explored in the paper, showcasing how specific, hidden behaviors can be embedded and subsequently detected within large language models.
Key Characteristics
- Backdoor Implementation: Contains an intentionally embedded "Harry Potter" themed backdoor, designed for research into hidden behaviors.
- Research Focus: Serves as a practical example for studying LLM interpretability, security, and the detection of covert functionalities.
- Architecture: Based on the AMF (Activation-matched Finetuning) framework, indicating its origin from a specific finetuning methodology.
- Parameter Count: Features 7.6 billion parameters, providing a substantial model size for demonstrating complex behaviors.
- Context Length: Supports a context window of 32768 tokens, allowing for analysis of longer inputs in research scenarios.
Intended Use Cases
- Academic Research: Ideal for researchers and academics investigating LLM security, backdoor detection, and model interpretability.
- Methodology Validation: Can be used to validate and test new methods for identifying and mitigating hidden behaviors in LLMs.
- Educational Tool: Useful for demonstrating the principles of activation-matched finetuning and the potential for embedded functionalities in language models.