typhoon-ai/llama3.1-typhoon2-deepseek-r1-70b-preview
Typhoon-AI's llama3.1-typhoon2-deepseek-r1-70b-preview is a 70 billion parameter Llama-based instruction-tuned model, specifically designed for enhanced reasoning in Thai and English. It merges DeepSeek R1 70B Distill and Typhoon2 70B Instruct + SFT, offering significantly improved math and coding performance compared to its predecessor. This model excels in cross-domain reasoning and maintains strong Thai language proficiency, making it suitable for complex analytical tasks.
Loading preview...
Typhoon2-DeepSeek-R1-70B: Enhanced Thai Reasoning LLM
Typhoon-AI's llama3.1-typhoon2-deepseek-r1-70b-preview is a 70 billion parameter large language model built on the Llama architecture, specifically engineered for advanced reasoning capabilities in both Thai and English. This research preview model is a merge of DeepSeek R1 70B Distill and Typhoon2 70B Instruct + SFT, aiming to combine their strengths.
Key Capabilities
- Advanced Reasoning: Offers reasoning capabilities comparable to DeepSeek R1 70B Distill.
- Enhanced Math & Coding: Demonstrates up to 6 times more accuracy in mathematical and coding tasks compared to Typhoon2 70B Instruct.
- Strong Thai Language Proficiency: Maintains high performance in Thai language understanding and generation.
- Cross-Domain Adaptability: Capable of adapting its reasoning skills across various domains, extending beyond just math and coding.
Good For
- Applications requiring robust reasoning in Thai and English.
- Tasks involving complex mathematical problems or code generation.
- Use cases where strong instruction-following and analytical capabilities are crucial.
This model is an instructional reasoning model and is still under development. Developers should assess potential risks related to accuracy or bias for their specific use cases. It utilizes the DeepSeek R1 70B Distill chat template, not the Llama 3 template, which is important for proper implementation.