FlagRelease/ERNIE-4.5-0.3B-PT-mthreads-FlagOS

TEXT GENERATIONPricing:Input $0.04 / Cached $0.002 / Output $0.08Concurrent Unit Cost:1Model Size:0.3BQuant:BF16Context Size:32kPublished:Jul 6, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

ERNIE-4.5-0.3B-PT-mthreads-FlagOS is a 0.3 billion parameter lightweight foundational LLM from Baidu Wenxin, specifically adapted for Moore Threads MTT S5000 GPUs. It leverages the FlagOS software stack, including MUSA & FlagGems acceleration, to provide optimized vLLM inference with significantly higher throughput on MUSA hardware. This model is designed for efficient deployment and performance on specific AI accelerators, offering out-of-the-box inference capabilities.

Loading preview...

ERNIE-4.5-0.3B-PT-mthreads-FlagOS Overview

This model is a 0.3 billion parameter foundational LLM from Baidu Wenxin, specifically engineered for optimal performance on Moore Threads MTT S5000 GPUs. It integrates with the FlagOS software stack, which unifies the "model–system–chip" layers to enable a "develop once, run anywhere" workflow across diverse AI accelerators. The model is provided with precompiled Triton cache and FlagGems operator acceleration.

Key Capabilities & Features

  • Hardware Optimization: Fully adapted for Moore Threads MTT S5000 GPUs using MUSA & FlagGems acceleration.
  • High-Throughput Inference: Supports out-of-the-box vLLM inference via the MUSA vLLM plugin (FlagTree backend), delivering significantly higher throughput compared to vanilla PyTorch on MUSA hardware.
  • Integrated Deployment: Comes with pre-configured hardware and software parameters, and a FlagOS-Mthreads container image for rapid deployment.
  • Consistency Validation: Rigorously evaluated against native stacks on public benchmarks, showing comparable performance to the NVIDIA-origin version.
  • FlagOS Ecosystem: Leverages core FlagOS technologies like FlagGems (high-performance operator library), FlagTree (unified compiler), and FlagScale (large model lifecycle toolkit).

Good For

  • Developers targeting Moore Threads MTT S5000 GPUs for LLM inference.
  • Use cases requiring high-throughput and efficient deployment of lightweight LLMs on specific accelerator hardware.
  • Environments benefiting from a unified software stack that simplifies model migration and deployment across different AI chips.