andrijdavid/Meta-Llama-3-13B-Instruct

TEXT GENERATIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:15BQuant:FP8Context Size:8kTool Calling:SupportedPublished:May 7, 2024License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

andrijdavid/Meta-Llama-3-13B-Instruct is a 15 billion parameter instruction-tuned causal language model, created by andrijdavid through a MergeKit self-merge of Meta-Llama-3-8B-Instruct. This model leverages a unique layer-slicing configuration to enhance capabilities, offering an 8192-token context length. It is designed for general-purpose conversational AI and instruction following tasks, building upon the robust foundation of the Llama 3 architecture.

Loading preview...

Overview

andrijdavid/Meta-Llama-3-13B-Instruct is a 15 billion parameter instruction-tuned language model. It was created by andrijdavid using MergeKit through a "self-merge" process, combining different layer ranges of the base Meta-Llama-3-8B-Instruct model. This merging technique aims to create a distinct model with potentially enhanced or altered characteristics compared to its source.

Key Characteristics

  • Architecture: Based on the Meta-Llama-3 family, known for strong performance in instruction following.
  • Parameter Count: 15 billion parameters, offering a balance between capability and computational requirements.
  • Context Length: Supports an 8192-token context window, suitable for handling moderately long inputs and generating coherent responses.
  • Merge Configuration: The model was constructed by merging specific layer ranges (0-16, 4-24, 8-31) of the 8B Instruct model, suggesting an attempt to combine different aspects or strengths of the original model's layers.

When to Use This Model

This model is suitable for developers looking for an instruction-tuned LLM derived from the Llama 3 architecture, particularly if they are interested in exploring models created via advanced merging techniques. Its 15B parameter count and 8K context window make it a strong candidate for various conversational AI applications, content generation, and instruction-following tasks where the specific merge configuration might offer unique performance characteristics.