EldritchLabs/MN-Aura-12B-v1
EldritchLabs/MN-Aura-12B-v1 is a 12.2 billion parameter language model created by EldritchLabs, merged using the Adaptive Unified Riemannian Annealing (AURA) method. This model integrates several pre-trained models, including Mistral-Nemo-Base-2407-ChatML as its base, to combine their strengths. It is designed for general language tasks, leveraging its merged architecture for broad applicability.
Loading preview...
Overview
EldritchLabs/MN-Aura-12B-v1 is a 12.2 billion parameter language model developed by EldritchLabs. It was constructed using the novel Adaptive Unified Riemannian Annealing (AURA) merge method, which is part of the mergekit-exp framework. The model's foundation is Retreatcost/Mistral-Nemo-Base-2407-ChatML, and it incorporates contributions from five other distinct 12B models, including DreadPoor/Famino-12B-Model_Stock, EldritchLabs/MN-Starlight-Sylph-12B, OccultAI/MN-Nazgul-12B-v1, shrugging-shoulders/Amberlight-Lux-12B, and WokeAI/Tankie-DPE-12B-SFT-v2. This merging strategy aims to synthesize the capabilities of multiple specialized models into a single, more versatile entity.
Key Capabilities
- Merged Architecture: Leverages the AURA merge method to combine the strengths of six different 12B parameter models.
- Base Model: Built upon the Retreatcost/Mistral-Nemo-Base-2407-ChatML, providing a strong foundation for chat-oriented tasks.
- Parameter Count: Operates with 12.2 billion parameters, offering a balance between performance and computational requirements.
- Context Length: Supports a context window of 32768 tokens, enabling processing of longer inputs and maintaining conversational coherence over extended interactions.
Good For
- General Language Tasks: Suitable for a wide range of applications due to its diverse merged components.
- Exploration of Merged Models: Ideal for developers interested in experimenting with models created via advanced merging techniques like AURA.
- Applications Requiring Moderate Scale: Provides a capable solution for use cases that benefit from a 12B parameter model with a substantial context window.