FlagRelease/Qwen3.8-27B-BF16-zhenwu-FlagOS-Express
FlagRelease/Qwen3.8-27B-BF16-zhenwu-FlagOS-Express is a 27 billion parameter vision-language model, derived from Alibaba's Qwen3.8-27B, optimized for multi-chip deployment. Developed by the Zhongzhi (众智) FlagOS community, this model features Day 0 multi-chip adaptation and precision alignment across 11 diverse AI chips, including NVIDIA and Ascend. It excels in providing out-of-the-box inference solutions with significant performance speedups, making it ideal for efficient, hardware-agnostic AI application deployment.
Loading preview...
Overview
FlagRelease/Qwen3.8-27B-BF16-zhenwu-FlagOS-Express is a 27 billion parameter vision-language model, based on Alibaba's Qwen3.8-27B, developed by the Zhongzhi (众智) FlagOS community. Its primary distinction is the Day 0 multi-chip adaptation and precision alignment across 11 different AI chips, including T-Head, NVIDIA, Moore Threads, and Ascend. This model is integrated into the FlagOS unified open-source technology stack, which aims to provide a "develop once, run anywhere" workflow for AI accelerators.
Key Capabilities
- Broad Hardware Compatibility: Adapted and verified across 11 diverse AI chips, with support for BF16 precision on most, and FP8 on NVIDIA and Moore Threads.
- Optimized Deployment: Offers out-of-the-box inference scripts and a FlagOS-Zhenwu container image for rapid deployment, enabling setup within minutes.
- Performance Enhancement: Demonstrates significant speedup ratios, with up to 188.68% improvement in certain concurrent test scenarios compared to native stacks.
- Edge-side Support: Extends adaptation to ARM edge-side platforms with a W4A8 low-bit version for broader accessibility.
- Robust Software Stack: Leverages FlagOS core technologies like FlagScale, FlagGems (high-performance operator library), FlagTree (unified compiler), and FlagCX (cross-chip communication library) for efficient model migration and hardware performance unlocking.
Good For
- Developers seeking a vision-language model with broad compatibility across various AI hardware platforms.
- Applications requiring efficient, optimized deployment and inference on diverse chip architectures.
- Users looking for out-of-the-box solutions to reduce the complexity and cost of porting and maintaining AI workloads across different hardware.