rinna/qwen2.5-bakeneko-32b

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:2Model Size:32.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Feb 10, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

rinna/qwen2.5-bakeneko-32b is a 32 billion parameter language model developed by rinna, continually pre-trained on 18 billion tokens of mixed Japanese and English datasets. This model is based on the Qwen2.5 architecture and is specifically optimized to enhance performance on Japanese language tasks. It serves as a strong foundation for applications requiring robust Japanese language understanding and generation.

Loading preview...

Overview

rinna/qwen2.5-bakeneko-32b is a 32 billion parameter language model developed by rinna, built upon the Qwen/Qwen2.5-32B architecture. It underwent continual pre-training on approximately 18 billion tokens, comprising a mixture of Japanese and English datasets, including Japanese CC-100, C4, OSCAR, The Pile, Wikipedia, and rinna's curated Japanese dataset. This extensive training significantly improves the model's proficiency in Japanese language tasks.

Key Capabilities

  • Enhanced Japanese Performance: Specialized continual pre-training on a large Japanese dataset improves its understanding and generation capabilities for Japanese text.
  • Foundation Model: Serves as a strong base for further fine-tuning, with instruction-tuned and reasoning-merged variants also available from rinna.
  • Qwen2.5 Architecture: Leverages the robust 64-layer, 5120-hidden-size transformer architecture of Qwen2.5.

Benchmarking

While the base rinna/qwen2.5-bakeneko-32b model shows a slight decrease in Japanese LM Evaluation Harness score compared to the original Qwen2.5-32B, its instruction-tuned variants, such as rinna/qwen2.5-bakeneko-32b-instruct and rinna/qwen2.5-bakeneko-32b-instruct-v2, demonstrate improved performance on Japanese MT-Bench scores, indicating better instruction following and conversational abilities in Japanese. For detailed results, refer to rinna's LM benchmark page.

Good for

  • Developers building applications focused on Japanese language processing.
  • Research and development requiring a strong Japanese-centric base model.
  • Further fine-tuning for specific Japanese NLP tasks.