MaziyarPanahi/japanese-stablelm-base-gamma-7b-Mistral-7B-Instruct-v0.2-slerp

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kTool Calling:SupportedPublished:Jan 12, 2024License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

MaziyarPanahi/japanese-stablelm-base-gamma-7b-Mistral-7B-Instruct-v0.2-slerp is a 7 billion parameter language model created by MaziyarPanahi, merging Mistral-7B-Instruct-v0.2 and japanese-stablelm-base-gamma-7b using the slerp method. This model combines the strong instruction-following capabilities of Mistral with the Japanese language proficiency of StableLM. It is designed for applications requiring robust performance in both English and Japanese contexts, leveraging a 4096-token context window.

Loading preview...

Model Overview

This model, japanese-stablelm-base-gamma-7b-Mistral-7B-Instruct-v0.2-slerp, is a 7 billion parameter language model developed by MaziyarPanahi. It is a product of merging two distinct base models: mistralai/Mistral-7B-Instruct-v0.2 and stabilityai/japanese-stablelm-base-gamma-7b. The merge was performed using the slerp (spherical linear interpolation) method, aiming to combine their respective strengths.

Key Characteristics

  • Hybrid Architecture: Integrates the instruction-tuned capabilities of Mistral-7B-Instruct-v0.2 with the Japanese language understanding of StableLM-Gamma-7B.
  • Parameter Count: Features 7 billion parameters, offering a balance between performance and computational efficiency.
  • Context Window: Supports a context length of 4096 tokens, suitable for various conversational and text generation tasks.
  • Merging Strategy: Utilizes a specific slerp configuration, applying different interpolation values to self-attention and MLP layers, indicating a fine-tuned approach to combining model weights.

Ideal Use Cases

  • Bilingual Applications: Particularly well-suited for scenarios requiring generation and understanding in both English and Japanese.
  • Instruction Following: Benefits from the instruction-tuned base, making it effective for tasks requiring precise responses to prompts.
  • Research and Development: Provides a strong foundation for further fine-tuning or experimentation in multilingual LLM development.