ukisai/Swift-1.5-Qwen3.8-27b

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 16, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

Swift 1.5 Qwen3.8-27B is a 27 billion parameter language model developed by UkisAI, derived from Qwen3.8-27B. This model is optimized for reasoning efficiency, using 58.5% fewer "thinking tokens" while achieving a 0.35% higher score than its base model, resulting in a 1.95x speed-up on various tasks. It excels particularly in coding and agentic tasks, making it suitable for applications requiring efficient and accurate problem-solving.

Loading preview...

Overview

Swift 1.5 Qwen3.8-27B is a 27 billion parameter model developed by UkisAI, building upon the Qwen3.8-27B architecture. It significantly improves reasoning efficiency, reducing "thinking tokens" by 58.5% while slightly increasing overall accuracy by 0.35% compared to the base model, leading to a 1.95x speed-up. This efficiency is achieved through scaled-up post-training (RL and OPD) that penalizes pathological overthinking tokens without directly shortening reasoning length.

Key Capabilities

  • Enhanced Reasoning Efficiency: Achieves higher accuracy with substantially fewer computational steps, particularly noticeable in tasks requiring complex thought processes.
  • Improved Performance: Outperforms both the base Qwen3.8-27B and its predecessor, Swift 1.0, across various benchmarks.
  • Specialized for Coding and Agentic Tasks: Demonstrates strong improvements in LiveCodeBench v6 and Terminal-Bench 2.1, making it highly effective for code generation and autonomous agent applications.
  • Flexible Reasoning Effort: Maintains token savings and competitive accuracy across different reasoning_effort settings (xhigh, medium, low) of the Qwen3.8 framework.
  • Quantized Versions Available: Offers various quantized formats (GGUF, GSQ-RCO GGUF, AWQ INT4, AutoRound INT4, etc.) for optimized deployment on different hardware.

Good for

  • Applications requiring fast and accurate reasoning: Ideal for scenarios where quick, high-quality responses are critical.
  • Code generation and development: Its strong performance in coding benchmarks makes it suitable for assisting developers.
  • Autonomous agent systems: Excels in agentic tasks, supporting the development of more efficient and capable AI agents.
  • Resource-constrained environments: The significant reduction in thinking tokens translates to more efficient resource utilization, especially with its diverse quantized options.