UNIVA-Bllossom/DeepSeek-qwen-Bllossom-32B

TEXT GENERATIONPricing:Input $1.06 / Cached $0.053 / Output $2.6Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 7, 2025License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

DeepSeek-qwen-Bllossom-32B is a 32 billion parameter language model developed jointly by UNIVA and Bllossom, built upon the DeepSeek-R1-Distill-Qwen-32B base. This model is specifically enhanced for Korean language inference, addressing performance degradation issues found in its English and Chinese-centric base model. It achieves improved reasoning capabilities in Korean environments by performing internal thought processes in English and generating responses in the input language. The model is optimized for accurate and reliable reasoning in Korean contexts, incorporating diverse reasoning data beyond STEM fields.

Loading preview...

DeepSeek-qwen-Bllossom-32B: Enhanced Korean Reasoning

DeepSeek-qwen-Bllossom-32B is a 32 billion parameter model developed through a collaboration between UNIVA and Bllossom. It is built upon the DeepSeek-R1-Distill-Qwen-32B base model, with a primary focus on significantly improving inference performance in Korean language environments. The base model, originally trained predominantly on English and Chinese data, exhibited performance drops when processing Korean.

Key Enhancements and Capabilities

  • Korean Language Optimization: Addresses the performance limitations of its base model in Korean by undergoing additional training with Korean and English reasoning datasets.
  • Internal Reasoning Strategy: The model is designed to conduct its internal thought processes in English, then generate responses in the user's input language, which enhances Korean output quality.
  • Diverse Training Data: Training incorporated a variety of reasoning data beyond the STEM fields typically used for DeepSeek-R1 models, aiming for broader applicability.
  • Post-training Distillation: Utilizes a post-training process with custom reasoning data to distill advanced reasoning and Korean processing capabilities into the DeepSeek-R1-Distill-Qwen-32B model, optimizing it for complex inference tasks.

Performance Benchmarks

The model demonstrates notable improvements in Korean benchmarks compared to its base model. For instance, on the AIME24_ko benchmark, DeepSeek-qwen-Bllossom-32B scores 66.67, significantly higher than DeepSeek-R1-Distill-Qwen-32B's 48.89. It also shows competitive performance on MATH500_ko and English benchmarks.

Licensing

DeepSeek-qwen-Bllossom-32B is licensed under the MIT License, allowing for commercial use, modifications, and derivative works, including distillation for other LLMs. Its base model, DeepSeek-R1-Distill-Qwen-32B, is derived from Qwen2.5-32B and is originally licensed under Apache 2.0.