AMAImedia/Qwen3.8-27B-Uncensored-NOESIS-BF16

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 4, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

AMAImedia/Qwen3.8-27B-Uncensored-NOESIS-BF16 is a 27 billion parameter Qwen3.8-based vision-language model developed by AMAImedia, derived from orcarouter's Qwen3.8-27B-Uncensored. This BF16 precision model features a hybrid Gated DeltaNet attention architecture, a native vision-language tower, and an MTP speculative-decoding head, with a context length of 262,144 tokens. Its primary differentiator is the surgical removal of safety alignment via 'abliteration,' making it suitable for research into refusal mechanisms, interpretability, red-teaming, and as a full-precision base for further fine-tuning and re-quantization.

Loading preview...

Model Overview

This model, AMAImedia/Qwen3.8-27B-Uncensored-NOESIS-BF16, is a 27 billion parameter, BF16 precision, vision-language model based on the Qwen3.8 architecture. Developed by AMAImedia as part of the NOESIS Professional Multilingual Dubbing Automation Platform, it is a modified version of orcarouter's Qwen3.8-27B-Uncensored. A key feature is its abliteration, a process that surgically removes safety alignment and refusal mechanisms, allowing it to comply with requests that the original model would refuse. It retains the full vision-language tower and MTP speculative-decoding head, supporting a substantial context length of 262,144 tokens.

Key Capabilities

  • Uncensored Responses: Safety alignment has been substantially removed, leading to compliance with harmful or unethical requests for research purposes.
  • Multimodal: Preserves the full vision-language tower, enabling image understanding and vision-conditioned responses.
  • High Precision: Provided in BF16 full precision, making it an ideal source for further fine-tuning, post-training (SFT, DPO, RLHF), and re-quantization to other formats like FP8 or GGUF.
  • Architectural Features: Utilizes a hybrid Gated DeltaNet attention (48 linear + 16 full attention layers) and an MTP speculative-decoding head.
  • Capability Retention: Evaluation shows essential capabilities like MMLU, GSM8K, and CMMLU are retained within a narrow margin of the base model, despite the abliteration.

Good For

  • Research: Ideal for studying refusal mechanisms, AI safety, interpretability, and red-teaming in controlled environments.
  • Fine-tuning & Post-training: Recommended as a full-precision base for custom fine-tuning, as it maintains the complete VL tower and MTP head.
  • Re-quantization: Suitable for generating quantized releases (e.g., FP8, GGUF) for deployment in environments with VRAM constraints.