ukisai/Swift-Qwen3.8-27b-BF16-AMD
Swift-Qwen3.8-27b-BF16-AMD is a 27 billion parameter BF16 model developed by UkisAI, derived from Qwen3.8-27B. This model is optimized for reasoning efficiency, achieving a x1.95 speed-up by using 58.3% fewer 'thinking tokens' with less than 1% performance loss. It is specifically designed as a full-precision companion for AMD/ROCm and AMD Quark workflows, including a vision tower and MTP head.
Loading preview...
Overview
Swift-Qwen3.8-27b-BF16-AMD is a 27 billion parameter model from UkisAI, built upon the Qwen3.8-27B architecture. Its core innovation lies in its reasoning efficiency, significantly reducing the number of 'thinking tokens' required for tasks. This results in a x1.95 speed-up on various benchmarks while maintaining near-identical performance (less than 1% loss compared to the base Qwen3.8-27B).
This specific release serves as the full-precision BF16 companion for AMD/ROCm and AMD Quark workflows, providing the merged Swift weights in standard Hugging Face safetensors format. It includes a vision tower and all 15 MTP (Multi-Task Prediction) tensors, making it suitable for multimodal applications and self-speculative decoding.
Key Capabilities & Differentiators
- Reasoning Efficiency: Achieves substantial reductions in 'thinking tokens' (e.g., 58.3% median reduction on GPQA-Diamond) leading to faster inference.
- Performance Retention: Maintains strong performance across general reasoning, mathematics, multimodal, and agentic coding benchmarks despite token reduction.
- AMD Optimization: Designed as a BF16 reference for AMD Quark INT4 quantization, supporting ROCm-compatible PyTorch/runtime builds.
- Multimodal Support: Includes a vision tower and MTP head for advanced capabilities.
- Quantization Compatibility: Demonstrates token savings and improved accuracy in quantized (INT4) deployments, particularly on tasks like AIME 2026.
When to Use This Model
- Resource-Constrained Environments: Ideal for deployments where faster inference and reduced token usage are critical, especially on AMD hardware.
- Reasoning-Intensive Tasks: Suitable for applications requiring complex reasoning, mathematics, and agentic coding, where efficiency is paramount.
- Multimodal Applications: Leverage its integrated vision tower for tasks involving both text and image inputs.
- Quantized Deployments: Excellent choice for scenarios requiring INT4 quantization, offering robust performance and token savings.