FlagRelease/Qwen3.8-27B-BF16-iluvatar-FlagOS

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

FlagRelease/Qwen3.8-27B-BF16-iluvatar-FlagOS is a 27 billion parameter vision-language model, derived from Alibaba's Qwen3.8-27B, specifically adapted by the Zhongzhi FlagOS community for multi-chip deployment. This model leverages the FlagOS unified open-source technology stack to achieve precision alignment and deployment verification across 11 different AI chips, including Iluvatar CoreX, with BF16 precision. It is designed for efficient, out-of-the-box inference and deployment across diverse hardware platforms, aiming to reduce fragmentation and lower AI workload porting costs.

Loading preview...

Overview

FlagRelease/Qwen3.8-27B-BF16-iluvatar-FlagOS is a 27 billion parameter vision-language model, based on Alibaba's Qwen3.8-27B, that has been extensively adapted by the Zhongzhi FlagOS community for multi-chip environments. It utilizes the FlagOS unified open-source technology stack to enable a "develop once, run anywhere" workflow across a wide range of AI accelerators. This model specifically highlights its adaptation and deployment verification across 11 different AI chips, including Iluvatar CoreX, with most running on BF16 precision.

Key Capabilities & Features

  • Multi-Chip Adaptation: Achieves precision alignment and deployment verification across 11 AI chips (T-Head, NVIDIA, Moore Threads, MetaX, Kunlunxin, Ascend, Hygon, Iluvatar CoreX, Tsingmicro, Enflame, Sunrise).
  • Precision Support: Supports BF16 precision on most adapted chips, with NVIDIA and Moore Threads also supporting FP8.
  • Edge-Side Deployment: Offers a W4A8 low-bit version for ARM edge-side platforms, providing out-of-the-box solutions.
  • Integrated Deployment: Provides out-of-the-box inference scripts and a dedicated FlagOS-Iluvatar container image for quick deployment.
  • Consistency Validation: Rigorously evaluated through benchmark testing to ensure performance and results consistency against native stacks.

FlagOS Technical Stack

This model's unique capabilities are powered by the FlagOS stack, which includes:

  • FlagScale: A distributed training/inference framework built on projects like Megatron-LM and vLLM, extended with vllm-plugin-fl for multi-chip support.
  • FlagGems: A high-performance, generic operator library implemented in Triton for accelerating LLM training and inference across diverse hardware.
  • FlagTree: An open-source, unified compiler for multiple AI chips, aiming to unify the codebase for multi-backend support.
  • FlagCX: A scalable and adaptive cross-chip communication library.

Evaluation Results

Benchmark results demonstrate competitive performance, for example:

  • musr_murder_mysteries: 77.56
  • GPQA_Diamond: 87.37

Good For

  • Developers seeking a Qwen3.8-27B variant optimized for deployment on a wide array of AI accelerators, particularly those supported by FlagOS.
  • Use cases requiring efficient, hardware-agnostic inference of large language models.
  • Environments needing streamlined deployment and consistent performance across heterogeneous computing infrastructures.