ukisai/Swift-Qwen3.8-27b-BF16-AMD

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 14, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

Swift-Qwen3.8-27b-BF16-AMD is a 27 billion parameter BF16 model developed by UkisAI, derived from Qwen3.8-27B. This model is optimized for reasoning efficiency, achieving a x1.95 speed-up by using 58.3% fewer 'thinking tokens' with less than 1% performance loss. It is specifically designed as a full-precision companion for AMD/ROCm and AMD Quark workflows, including a vision tower and MTP head.

Loading preview...

Overview

Swift-Qwen3.8-27b-BF16-AMD is a 27 billion parameter model from UkisAI, built upon the Qwen3.8-27B architecture. Its core innovation lies in its reasoning efficiency, significantly reducing the number of 'thinking tokens' required for tasks. This results in a x1.95 speed-up on various benchmarks while maintaining near-identical performance (less than 1% loss compared to the base Qwen3.8-27B).

This specific release serves as the full-precision BF16 companion for AMD/ROCm and AMD Quark workflows, providing the merged Swift weights in standard Hugging Face safetensors format. It includes a vision tower and all 15 MTP (Multi-Task Prediction) tensors, making it suitable for multimodal applications and self-speculative decoding.

Key Capabilities & Differentiators

  • Reasoning Efficiency: Achieves substantial reductions in 'thinking tokens' (e.g., 58.3% median reduction on GPQA-Diamond) leading to faster inference.
  • Performance Retention: Maintains strong performance across general reasoning, mathematics, multimodal, and agentic coding benchmarks despite token reduction.
  • AMD Optimization: Designed as a BF16 reference for AMD Quark INT4 quantization, supporting ROCm-compatible PyTorch/runtime builds.
  • Multimodal Support: Includes a vision tower and MTP head for advanced capabilities.
  • Quantization Compatibility: Demonstrates token savings and improved accuracy in quantized (INT4) deployments, particularly on tasks like AIME 2026.

When to Use This Model

  • Resource-Constrained Environments: Ideal for deployments where faster inference and reduced token usage are critical, especially on AMD hardware.
  • Reasoning-Intensive Tasks: Suitable for applications requiring complex reasoning, mathematics, and agentic coding, where efficiency is paramount.
  • Multimodal Applications: Leverage its integrated vision tower for tasks involving both text and image inputs.
  • Quantized Deployments: Excellent choice for scenarios requiring INT4 quantization, offering robust performance and token savings.