renaudb1999/le-harnais-ft-counsel-Llama-3.2-3B-Instruct-jepa-full
The renaudb1999/le-harnais-ft-counsel-Llama-3.2-3B-Instruct-jepa-full is a 3.2 billion parameter Llama-3.2-3B-Instruct based model developed by renaudb1999. This specific model is an ablation checkpoint focused on studying counsel-corpus scaling, data augmentation, and JEPA (Joint Embedding Predictive Architecture) effects. It is designed for research into training methodologies rather than direct inference, serving as a component for reproducing and continuing training studies.
Loading preview...
Model Overview
This model, le-harnais-ft-counsel-Llama-3.2-3B-Instruct-jepa-full, is an ablation checkpoint derived from the meta-llama/Llama-3.2-3B-Instruct base model. It is specifically designed for research purposes, focusing on the impact of counsel-corpus scaling, data augmentation, and the Joint Embedding Predictive Architecture (JEPA) on model performance.
Key Characteristics
- Base Model: Built upon
meta-llama/Llama-3.2-3B-Instruct. - Parameter Count: 3.2 billion parameters.
- Context Length: 32768 tokens.
- Training Data: Utilizes
datasets/counsel_train.jsonl, comprising 270 examples of wisdom and commentary from public domain sources. - JEPA Integration: Explores the effect of JEPA, noting a gain of approximately +7 at 3B parameters, but around 0 at 8B.
Intended Use
This model is not intended for direct inference or general use cases. Its primary purpose is to facilitate the reproduction and continuation of training studies, particularly those investigating the counsel-corpus scaling grid and JEPA's influence. For practical inference, users are directed to the "hero models" such as le-harnais-ft-agentworld-{1b,3b,8b} or le-harnais-ft-counsel.