tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.1

Hugging Face
TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 24, 2024License:llama3.1Architecture:Transformer0.0K Featherless Exclusive Warm

The Llama-3.1-Swallow-8B-Instruct-v0.1 is an 8 billion parameter instruction-tuned large language model developed by tokyotech-llm, built upon Meta Llama 3.1. It features enhanced Japanese language capabilities through continual pre-training on a 200 billion token Japanese web corpus, while retaining strong English performance. This model excels in Japanese language tasks, making it suitable for applications requiring robust bilingual understanding and generation, with a context length of 32768 tokens.

Loading preview...

Llama 3.1 Swallow 8B Instruct v0.1: Enhanced Japanese Capabilities

This model is an 8 billion parameter instruction-tuned variant from the Llama 3.1 Swallow series, developed by tokyotech-llm. It is built by continually pre-training on the Meta Llama 3.1 base models, specifically focusing on significantly enhancing Japanese language capabilities while maintaining strong English performance.

Key Capabilities

  • Bilingual Proficiency: Excels in both Japanese and English, with a particular focus on Japanese language tasks.
  • Continual Pre-training: Utilizes approximately 200 billion tokens from a large Japanese web corpus (Swallow Corpus Version 2), Japanese and English Wikipedia, and mathematical/coding content.
  • Instruction-Tuned: Supervised fine-tuning (SFT) was performed using synthetic data specifically designed for Japanese.
  • Strong Japanese Benchmarks: Achieves leading scores in various Japanese evaluation benchmarks, including JCommonsenseQA, JEMHopQA, NIILC, and JSQuAD, and shows competitive performance in MT-Bench JA.
  • Llama 3.1 Foundation: Benefits from the robust architecture and tokenizer of the Meta Llama 3.1 models.

Good for

  • Applications requiring high-quality Japanese language understanding and generation.
  • Bilingual (Japanese-English) conversational AI and instruction-following tasks.
  • Research and development in cross-lingual LLM adaptation, particularly for Japanese.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p