KantaHayashiAI/mamba-char-japanese-790m
KantaHayashiAI/mamba-char-japanese-790m is an experimental 0.79 billion parameter Mamba-architecture language model developed by KantaHayashiAI, specifically designed for Japanese text. It utilizes the ku-nlp/gpt2-large-japanese-char tokenizer and is released under the Apache 2.0 license. This model is an early-stage project, not yet fully trained, focusing on exploring Mamba's capabilities for Japanese character-level processing.
Loading preview...
Model Overview
KantaHayashiAI/mamba-char-japanese-790m is an experimental Japanese language model built on the Mamba architecture, featuring 0.79 billion parameters. Developed by KantaHayashiAI, this model is in its early stages of training and is intended for research and exploration into Mamba's performance with Japanese text.
Key Characteristics
- Architecture: Mamba, a state-space model known for its efficiency and performance.
- Language: Primarily focused on Japanese, utilizing a character-level tokenizer.
- Tokenizer: Employs the
ku-nlp/gpt2-large-japanese-chartokenizer for processing Japanese text at a character granularity. - License: The model is distributed under the Apache 2.0 license, while its tokenizer uses CC-BY-SA.
- Context Length: Supports a context window of 32768 tokens.
- Development Status: This is an experimental model that has not yet undergone sufficient training, indicating it is a work in progress.
Intended Use
This model is suitable for:
- Researchers and developers interested in experimenting with Mamba architecture for Japanese language tasks.
- Exploring character-level processing in Japanese LLMs.
- Early-stage prototyping and understanding the behavior of Mamba models before extensive training.
It is important to note that due to its experimental and insufficiently trained status, it is not recommended for production environments or applications requiring high accuracy and robustness.