renaudb1999/le-harnais-ft-counsel-Meta-Llama-3.1-8B-Instruct-jepa-full
This model, renaudb1999/le-harnais-ft-counsel-Meta-Llama-3.1-8B-Instruct-jepa-full, is an 8 billion parameter ablation checkpoint based on Meta-Llama-3.1-8B-Instruct. It is specifically designed for research into counsel-corpus scaling, data augmentation, and JEPA (Joint Embedding Predictive Architecture) effects. This version is intended for reproducing or continuing training studies and is not recommended for general inference due to its experimental nature.
Loading preview...
Model Overview
This model, le-harnais-ft-counsel-Meta-Llama-3.1-8B-Instruct-jepa-full, is an 8 billion parameter ablation checkpoint derived from the meta-llama/Meta-Llama-3.1-8B-Instruct base model. It is part of the le-harnais project, focusing on experimental research rather than direct inference applications.
Key Characteristics
- Base Model: Built upon
meta-llama/Meta-Llama-3.1-8B-Instruct, adhering to the Llama Community License. - Purpose: Classified as an
ablationmodel, its primary role is to facilitate the reproduction and continuation of training studies. - Training Focus: Explores counsel-corpus scaling, data augmentation techniques, and the impact of JEPA (Joint Embedding Predictive Architecture).
- JEPA Ablation: Specifically investigates JEPA's effect, noting an approximate +7 improvement at 3B parameters but negligible impact at 8B parameters within this experimental context.
- Training Data: Utilizes
datasets/counsel_train.jsonl, comprising 270 examples of wisdom and commentary from public domain sources.
Important Note
This checkpoint is explicitly marked as "retrain-only" and "not useful for inference." Developers should use this model exclusively for research, reproduction, or continuation of the specified training study. For general inference, users are directed to the project's hero models such as le-harnais-ft-agentworld-{1b,3b,8b} or le-harnais-ft-counsel.