grimjim/Magnolia-v3-medis-remix-12B
grimjim/Magnolia-v3-medis-remix-12B is a 12 billion parameter language model, merged using the Task Arithmetic method with a Mistral Nemo base and a 32768 token context length. It incorporates several pre-trained models, including a medical fine-tune, to enhance its capabilities. This model is designed to leverage the Mistral Nemo architecture, optimized for instruction-following tasks using the Tekken Instruct Chat Template. Its unique composition suggests a focus on diverse applications, potentially including specialized domains due to the medical component.
Loading preview...
Model Overview
Magnolia-v3-medis-remix-12B is a 12 billion parameter language model developed by grimjim, created through a sophisticated merge of several pre-trained models using the Task Arithmetic method. Built upon a grimjim/mistralai-Mistral-Nemo-Base-2407 foundation, this model integrates components from grimjim/magnum-consolidatum-v1-12b, exafluence/EXF-Medistral-Nemo-12B, nbeerbower/Mistral-Nemo-Prism-12B, and grimjim/magnum-twilight-12b, alongside grimjim/mistralai-Mistral-Nemo-Instruct-2407.
Key Capabilities
- Merged Architecture: Leverages the strengths of multiple models, including a significant contribution from Nemo Instruct.
- Medical Component: Incorporates a medical fine-tune (
exafluence/EXF-Medistral-Nemo-12B) as a "noise" component, suggesting potential for specialized domain understanding. - Instruction Following: Tuned to work with Mistral's Tekken Instruct Chat Template, similar to their Tokenizer V3, for effective instruction-based interactions. More details on the chat template can be found in Mistral's documentation.
- Context Length: Supports a context length of 32768 tokens.
Good For
- Applications requiring a blend of general instruction-following and potentially specialized domain knowledge, particularly in areas where the medical fine-tune might be beneficial.
- Developers familiar with Mistral's chat templates and tokenization schemes.