LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Final-Safetensors

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 3, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Final-Safetensors is a 35.1 billion parameter Qwen3.6 MoE model, reconstructed in SafeTensors/BF16 format from a Q8_K_P GGUF lineage. It features approximately 3 billion active parameters per forward pass, a 32768-token context length, and native multimodal capabilities including text, image, and video processing. This model is optimized for Transformers/vLLM inference on NVIDIA GPUs, particularly for precise tasks like coding and reasoning, and supports tool calling.

Loading preview...

Qwen3.6-35B-A3B-Uncensored-Genesis-Final-Safetensors Overview

This model is a SafeTensors/BF16 reconstruction of the LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Final-GGUF model, primarily intended for Transformers / vLLM inference on NVIDIA GPUs. It was reconstructed from a Q8_K_P GGUF lineage, meaning it is not the original publisher's BF16 checkpoint, and information lost during quantization cannot be recovered. Architecture-required tensors not present in the GGUF were copied from the compatible Qwen/Qwen3.6-35B-A3B reference model.

Key Capabilities

  • MoE Architecture: Features a Qwen3.6 MoE architecture with ~35B total parameters and ~3B active parameters per forward pass, utilizing 256 experts with 8 routed + 1 shared expert per token.
  • Multimodal Support: Native multimodal capabilities for text, image, and video processing, with the Hugging Face multimodal processor and vision tensors integrated directly into the model layout.
  • Extended Context: Supports a native context metadata up to 262K tokens, with validated vLLM configurations supporting 131072 tokens.
  • Tool Calling & Reasoning: Configurable for Qwen-style reasoning parsing and automatic tool choice, compatible with OpenAI-compatible clients like OpenCode.
  • Optimized for NVIDIA GPUs: Verified for performance on NVIDIA RTX PRO 6000 Blackwell Workstation Edition, with specific vLLM setup recommendations for optimal performance.

Good for

  • High-Performance Inference: Ideal for developers requiring a Qwen3.6-based model for vLLM inference on NVIDIA GPUs, especially Blackwell architecture.
  • Precise Tasks & Coding: Recommended sampling settings are provided for "thinking mode" tasks, including coding and other precise applications.
  • Multimodal Applications: Suitable for applications requiring integrated text, image, and video processing capabilities within a single model.
  • Tool-Augmented Workflows: Excellent for use cases involving tool calling and structured reasoning, leveraging its Qwen-style parsing capabilities.