theocolf/vora-x-bloc1
Theocolf/vora-x-bloc1 is an 8 billion parameter language model created by theocolf, merged from unsloth/llama-3-8b-Instruct, cognitivecomputations/dolphin-2.9-llama3-8b, and HumanLLMs/Human-Like-LLama3-8B-Instruct using a linear merge method. This model leverages the strengths of its constituent Llama 3-based models, offering a balanced performance profile for general language tasks. It maintains an 8192-token context length, suitable for a variety of conversational and text generation applications.
Loading preview...
Model Overview
The theocolf/vora-x-bloc1 is an 8 billion parameter language model developed by theocolf, created through a linear merge of several pre-trained models. This merge process combines the characteristics of its base models to offer a versatile language understanding and generation capability.
Merge Details
This model was constructed using the MergeKit tool with a linear merge method. The primary base model for this merge was unsloth/llama-3-8b-Instruct.
Constituent Models
The vora-x-bloc1 integrates components from the following models:
- unsloth/llama-3-8b-Instruct: Served as the foundational base model.
- cognitivecomputations/dolphin-2.9-llama3-8b: Contributes to the model's overall performance.
- HumanLLMs/Human-Like-LLama3-8B-Instruct: Further enhances the merged model's capabilities.
Each model was assigned specific weights during the linear merge process (0.34 for unsloth/llama-3-8b-Instruct, 0.33 for cognitivecomputations/dolphin-2.9-llama3-8b, and 0.33 for HumanLLMs/Human-Like-LLama3-8B-Instruct), aiming to balance their respective strengths. The tokenizer from the base model was copied to ensure compatibility and consistent tokenization.
Key Characteristics
- Parameter Count: 8 billion parameters.
- Context Length: Supports an 8192-token context window.
- Architecture: Based on the Llama 3 family, inheriting its robust architecture.
- Development Method: Created via a linear merge, combining multiple instruction-tuned Llama 3 variants.
Potential Use Cases
This model is suitable for a range of applications where a balanced performance from Llama 3-based models is desired, including:
- General-purpose text generation.
- Conversational AI and chatbots.
- Instruction-following tasks.
- Text summarization and analysis.