FlagRelease/Qwen3.8-27B-BF16-nvidia-FlagOS

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

FlagRelease/Qwen3.8-27B-BF16-nvidia-FlagOS is a 27 billion parameter vision-language model, derived from Alibaba's Qwen3.8-27B, specifically adapted and optimized for NVIDIA hardware within the FlagOS ecosystem. This model is part of a multi-chip adaptation initiative by the Zhongzhi FlagOS community, ensuring deployment verification and precision alignment across various AI chips. It is designed for efficient, out-of-the-box inference on NVIDIA platforms, leveraging the FlagOS unified open-source technology stack for streamlined AI workload deployment.

Loading preview...

Overview

FlagRelease/Qwen3.8-27B-BF16-nvidia-FlagOS is a 27 billion parameter vision-language model, originating from Alibaba's Qwen3.8-27B. This specific release is optimized for NVIDIA hardware, showcasing the Zhongzhi FlagOS community's efforts in multi-chip adaptation and deployment verification across 11 different AI chips. It supports BF16 precision, with FP8 precision available for NVIDIA and Moore Threads.

Key Capabilities & Features

  • Multi-Chip Adaptation: Part of a broader initiative to adapt Qwen models across diverse AI accelerators, including NVIDIA, T-Head, Moore Threads, and others.
  • FlagOS Integration: Leverages the FlagOS unified open-source technology stack, which includes FlagScale, FlagGems, FlagCX, and FlagTree, to enable a "develop once, run anywhere" workflow.
  • Out-of-the-Box Deployment: Provides pre-configured hardware and software parameters with a dedicated FlagOS-Nvidia container image for rapid deployment.
  • Consistency Validation: Rigorously evaluated through benchmark testing, demonstrating performance consistency with the original Qwen3.8-27B model on NVIDIA.
  • Vision-Language Model: Inherits the vision-language capabilities of the base Qwen3.8-27B model.

Good For

  • Developers seeking an optimized Qwen3.8-27B model for NVIDIA GPUs.
  • Users requiring streamlined, out-of-the-box deployment solutions for large language models.
  • Environments focused on multi-chip compatibility and reducing AI workload porting costs through the FlagOS ecosystem.
  • Applications benefiting from a vision-language model with verified performance on NVIDIA hardware.