FlagRelease/Qwen3.8-27B-BF16-tsingmicro-FlagOS
FlagRelease/Qwen3.8-27B-BF16-tsingmicro-FlagOS is a 27 billion parameter vision-language model, derived from Alibaba's Qwen3.8-27B, optimized for multi-chip deployment. Developed by the FlagOS community, it features adaptation and precision alignment across 11 AI chips, including Tsingmicro, with BF16 precision. This model is specifically designed for efficient, out-of-the-box deployment on diverse hardware, extending to ARM edge-side platforms with W4A8 low-bit versions, making it suitable for broad AI application deployment.
Loading preview...
Overview
FlagRelease/Qwen3.8-27B-BF16-tsingmicro-FlagOS is a 27 billion parameter vision-language model, originating from Alibaba's Qwen3.8-27B, that has been extensively adapted and optimized by the FlagOS community for multi-chip deployment. It supports BF16 precision on most platforms and FP8 on NVIDIA and Moore Threads, with a notable extension to W4A8 low-bit versions for ARM edge-side platforms.
Key Capabilities & Features
- Broad Hardware Compatibility: Adapted for 11 different AI chips, including T-Head, NVIDIA, Moore Threads, MetaX, Kunlunxin, Ascend, Hygon, Iluvatar CoreX, Tsingmicro, Enflame, and Sunrise.
- FlagOS Integration: Leverages the FlagOS unified open-source technology stack, which includes FlagScale, FlagGems, FlagCX, and FlagTree, to enable a "develop once, run anywhere" workflow.
- Out-of-the-Box Deployment: Provides integrated deployment solutions, including a FlagOS-Tsingmicro container image for quick setup and inference scripts.
- Precision Alignment: Features precision alignment and deployment verification across diverse hardware, supporting BF16 and FP8 (on select chips) and W4A8 for edge devices.
Use Cases
This model is ideal for developers and organizations seeking a versatile and hardware-agnostic large language model. Its multi-chip adaptation and optimized deployment solutions make it particularly suitable for:
- Deploying LLM applications across heterogeneous AI hardware environments.
- Edge computing scenarios requiring low-bit precision (W4A8) on ARM platforms.
- Reducing the cost and complexity of porting and maintaining AI workloads across different vendor-specific software stacks.