jpquiroga/Mistral_7B_ties_merge_instruct_open_orca_codeninja
jpquiroga/Mistral_7B_ties_merge_instruct_open_orca_codeninja is a 7 billion parameter language model based on the Mistral-7B-v0.1 architecture, created by jpquiroga. This model is a TIES merge of Mistral-7B-Instruct-v0.1, Open-Orca/Mistral-7B-OpenOrca, and beowolx/CodeNinja-1.0-OpenChat-7B. It is designed to combine the strengths of instruction following, general reasoning, and code generation, making it suitable for diverse conversational and programming tasks.
Loading preview...
Model Overview
This model, jpquiroga/Mistral_7B_ties_merge_instruct_open_orca_codeninja, is a 7 billion parameter language model built upon the mistralai/Mistral-7B-v0.1 base. It was created by jpquiroga using the TIES merge method via mergekit.
Key Capabilities
This model integrates the strengths of three distinct Mistral-7B variants:
- Instruction Following: Incorporates
mistralai/Mistral-7B-Instruct-v0.1for robust instruction adherence and conversational abilities. - Reasoning and General Knowledge: Leverages
Open-Orca/Mistral-7B-OpenOrcato enhance its general reasoning and understanding capabilities. - Code Generation: Includes
beowolx/CodeNinja-1.0-OpenChat-7Bto provide strong performance in code-related tasks, making it proficient in generating and understanding programming constructs.
Merge Configuration
The TIES merge method was applied with specific density and weight parameters for each contributing model, ensuring a balanced integration of their respective features. The merge utilized bfloat16 for its data type and included int8_mask and normalization during the process. This strategic combination aims to produce a versatile model capable of handling a broad spectrum of natural language processing and code-centric applications.