unsloth/ERNIE-4.5-21B-A3B-Thinking
ERNIE-4.5-21B-A3B-Thinking is a 21 billion parameter text Mixture-of-Experts (MoE) model developed by Baidu, featuring 3 billion activated parameters per token and a 131,072 token context length. This model is specifically enhanced for complex reasoning tasks, demonstrating significantly improved performance across logical reasoning, mathematics, science, coding, and academic benchmarks. It also offers efficient tool usage capabilities and advanced long-context understanding, making it suitable for applications requiring deep analytical thought.
Loading preview...
ERNIE-4.5-21B-A3B-Thinking: Enhanced Reasoning MoE Model
ERNIE-4.5-21B-A3B-Thinking is a text-based Mixture-of-Experts (MoE) model from Baidu, featuring 21 billion total parameters with 3 billion activated parameters per token. It is designed to significantly advance reasoning capabilities, building upon the ERNIE-4.5-21B-A3B series.
Key Enhancements and Capabilities
- Superior Reasoning Performance: Demonstrates marked improvements in complex reasoning tasks, including logical reasoning, mathematical problem-solving, scientific inquiry, code generation, and general text generation. It excels in academic benchmarks that typically demand human-level expertise.
- Efficient Tool Usage: Equipped with enhanced capabilities for integrating and utilizing external tools, supporting more dynamic and functional interactions.
- Extended Context Understanding: Features an impressive 131,072 token context length, enabling deeper comprehension and processing of very long inputs.
- MoE Architecture: Utilizes a Mixture-of-Experts design with 64 total text experts (6 activated) and 64 total vision experts (6 activated), alongside 2 shared experts, contributing to its efficiency and performance.
Recommended Use Cases
This model is particularly recommended for highly complex reasoning tasks where deep analytical thought and extensive context understanding are critical. Its improved 'thinking length' makes it a strong candidate for applications requiring advanced problem-solving and intricate logical deductions.