ark4004/DeepSeek-llama3.3-Bllossom-70B-jh

TEXT GENERATIONPricing:Input $2.88 / Output $2.88Concurrent Unit Cost:4Model Size:70BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 12, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The ark4004/DeepSeek-llama3.3-Bllossom-70B-jh is a 70 billion parameter language model developed by UNIVA and Bllossom, built upon the DeepSeek-R1-distill-Llama-70B base model with a 32768 token context length. It is specifically enhanced to improve Korean language inference performance, addressing limitations of its English and Chinese-centric base model. This model achieves improved reasoning capabilities in Korean environments by performing internal thought processes in English and generating responses in the input language.

Loading preview...

DeepSeek-llama3.3-Bllossom-70B: Enhanced Korean Reasoning

DeepSeek-llama3.3-Bllossom-70B is a 70 billion parameter model developed by UNIVA and Bllossom, designed to significantly improve Korean language inference performance. Built on the DeepSeek-R1-distill-Llama-70B base model, it addresses the base model's limitations in multilingual contexts, particularly for Korean.

Key Capabilities & Enhancements

  • Improved Korean Inference: The model is specifically post-trained to enhance reasoning in Korean environments. It processes internal thoughts in English and then generates responses in the user's input language, leading to more accurate and reliable Korean outputs.
  • Diverse Training Data: Training included Korean and English reasoning datasets, expanding beyond the STEM-focused data of the original DeepSeek-R1 models to cover a wider range of domains.
  • Post-training for Reasoning: Utilizes a post-training process with custom reasoning data to distill advanced reasoning and Korean processing capabilities into the base model, optimizing it for complex inference tasks.

Performance Highlights

Benchmarks show improved performance on Korean-specific tasks compared to its base model. For instance, on the AIME24_ko benchmark, DeepSeek-llama3.3-Bllossom-70B scored 62.22, an increase from the DeepSeek-R1-Distill-Llama-70B's 58.89. Similarly, on MATH500_ko, it achieved 88.40.

Licensing

This model and its code repository are licensed under the MIT License, supporting commercial use and modifications, including distillation for other LLMs. It is derived from the Llama3.3-70B-Instruct and DeepSeek-R1-Distill-Llama-70B models, retaining their original Llama3.3 license.