tokyotech-llm/Gemma-2-Llama-Swallow-9b-pt-v0.1

Hugging Face
TEXT GENERATIONPricing:Input $0.431 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:16kPublished:Mar 9, 2025License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Warm

The tokyotech-llm/Gemma-2-Llama-Swallow-9b-pt-v0.1 is a 9 billion parameter pre-trained language model developed by tokyotech-llm, built upon the Gemma 2 architecture. This model enhances the Japanese language capabilities of the original Gemma 2 while retaining strong English performance, achieved through continual pre-training on approximately 200 billion tokens from Japanese web corpora, Wikipedia, and mathematical/coding content. It is designed for general language understanding and generation tasks, particularly excelling in Japanese contexts.

Loading preview...

Model Overview

The tokyotech-llm/Gemma-2-Llama-Swallow-9b-pt-v0.1 is a 9 billion parameter pre-trained language model developed by tokyotech-llm. It is part of the Gemma-2-Llama-Swallow series, which continually pre-trains on Google's Gemma 2 models to enhance Japanese language capabilities while maintaining English proficiency.

Key Capabilities & Training

  • Bilingual Proficiency: Significantly improved Japanese language understanding and generation, alongside strong English performance.
  • Extensive Pre-training: Continually pre-trained on approximately 200 billion tokens, including a large Japanese web corpus (Swallow Corpus Version 2), Japanese and English Wikipedia, and specialized mathematical and coding content.
  • Architecture: Based on the Gemma 2 model architecture, leveraging its foundational strengths.

Performance Highlights

This model demonstrates competitive performance across various benchmarks:

  • Japanese Tasks: Achieves a Ja Avg score of 0.558, outperforming google/gemma-2-9b (0.500) and google/gemma-3-12b-pt (0.518) in its size class.
  • English Tasks: Achieves an En Avg score of 0.595, comparable to google/gemma-2-9b (0.597).

Good for

  • Applications requiring strong bilingual (Japanese and English) language understanding and generation.
  • Research and development in Japanese NLP, leveraging a robust pre-trained base.
  • Tasks involving general text processing, question answering, and content generation in both languages.