unsloth/ERNIE-4.5-0.3B-PT

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.3BQuant:BF16Context Size:32kPublished:Jun 30, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The ERNIE-4.5-0.3B-PT model, developed by Baidu, is a 0.36 billion parameter text-dense post-trained language model with a 131,072 token context length. It is part of the ERNIE 4.5 series, which features multimodal heterogeneous MoE pre-training and scaling-efficient infrastructure. This specific variant is optimized for general-purpose language understanding and generation tasks, leveraging advanced post-training methods like Supervised Fine-tuning (SFT) and Unified Preference Optimization (UPO).

Loading preview...

ERNIE-4.5-0.3B-PT Model Overview

ERNIE-4.5-0.3B-PT is a 0.36 billion parameter text-dense language model developed by Baidu, featuring an extensive context length of 131,072 tokens. It is a PyTorch-based variant from the ERNIE 4.5 series, which is distinguished by its innovative multimodal heterogeneous Mixture-of-Experts (MoE) pre-training approach. This pre-training involves joint training on both textual and visual modalities, utilizing techniques like modality-isolated routing and multimodal token-balanced loss to ensure effective representation and mutual reinforcement across modalities.

Key Technical Innovations

  • Multimodal Heterogeneous MoE Pre-Training: Designed to capture nuances from both text and visual information, enhancing performance in cross-modal reasoning tasks. The architecture employs a heterogeneous MoE structure with modality-isolated routing and specific loss functions to prevent one modality from hindering another's learning.
  • Scaling-Efficient Infrastructure: Features a novel heterogeneous hybrid parallelism and hierarchical load balancing strategy for efficient training. This includes intra-node expert parallelism, memory-efficient pipeline scheduling, FP8 mixed-precision training, and fine-grained recomputation. For inference, it introduces multi-expert parallel collaboration and convolutional code quantization for 4-bit/2-bit lossless quantization.
  • Modality-Specific Post-Training: This particular 0.3B model is post-trained specifically for text, optimized for general-purpose language understanding and generation. It leverages Supervised Fine-tuning (SFT), Direct Preference Optimization (DPO), or Unified Preference Optimization (UPO) for refinement.

Use Cases

This model is well-suited for applications requiring robust text understanding and generation, especially where a compact model with a large context window is beneficial. Its advanced training methodologies make it a strong candidate for various natural language processing tasks.