Mantisec/Qwen3.8-27B-TURBO-FP16

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 13, 2026Architecture:Transformer0.0K Featherless Exclusive Cold

Mantisec/Qwen3.8-27B-TURBO-FP16 is an FP16 conversion of the DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU model, featuring 27 billion parameters and a 32,768 token context length. This version is specifically optimized for efficient inference and fine-tuning on NVIDIA V100 (Volta) GPUs, which lack native BF16 Tensor Core support. The conversion uses a range-checked strategy to preserve numerical quality, achieving a 99.80% token agreement rate with the source model. It is designed for applications requiring high performance on specific hardware configurations.

Loading preview...

Mantisec/Qwen3.8-27B-TURBO-FP16 Overview

Mantisec/Qwen3.8-27B-TURBO-FP16 is a 27 billion parameter language model, derived from DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU. This model has been meticulously converted to FP16 using the bfsquish v0.1.0 tool, specifically targeting optimal performance on NVIDIA V100 (Volta) GPUs.

Key Features & Optimizations

  • FP16 Conversion: Optimized for hardware lacking native BF16 support, such as NVIDIA V100 GPUs, enabling efficient inference and fine-tuning.
  • Range-Checked Strategy: The conversion process employs a range_checked strategy, ensuring high numerical fidelity by directly upcasting to FP32 and then performing a range-checked FP16 cast. This minimizes data loss and preserves the source checkpoint's characteristics.
  • High Numerical Quality: Validation metrics show a PASS verdict, with a token agreement rate of 99.80% and a minimum cosine similarity of 0.999265 compared to the source model, indicating robust preservation of model behavior.
  • Enhanced Chat Template: Includes an enhanced chat template from peculiar-ragdoll/Qwen-Sharp-Chat-Templates, providing improved prompt rendering capabilities without altering the model weights.
  • vLLM Warm-up Profile: A version 1 advisory profile is provided for vLLM, offering hints for text generation warm-up, though runtimes may override these suggestions.

Use Cases

This model is particularly well-suited for developers and researchers who:

  • Require a high-performance 27B parameter model for inference or fine-tuning on NVIDIA V100 (Volta) GPUs.
  • Need a numerically stable FP16 version of the Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU model.
  • Are working in environments where FP16 precision is preferred or necessary for computational efficiency.