ericpandev/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-bf16

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 8, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

ericpandev/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-bf16 is a bf16 safetensors conversion of the Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive model, originally created by HauhauCS. This multimodal Qwen3.6 35B MoE architecture features 256 experts (8 active, 3B active parameters) and a substantial context length of 262,144 tokens. It is designed for use with vLLM speculators as a verifier/target model, offering a full tokenizer with 248,320 vocabulary and a 27-layer vision encoder for multimodal applications.

Loading preview...

Model Overview

This model, ericpandev/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-bf16, is a bf16 safetensors conversion of the original GGUF model by HauhauCS. It is based on the Qwen 3.6 35B MoE architecture, featuring 256 experts with 8 active, totaling 3 billion active parameters. The conversion ensures compatibility with the HuggingFace format, making it ready for vLLM speculators (DFlash/EAGLE-3) as a verifier or target model.

Key Capabilities & Specifications

  • Architecture: Qwen3_5MoeForConditionalGeneration, indicating multimodal capabilities.
  • Expert Configuration: Utilizes 256 routed experts plus 1 shared expert per layer, with 8 active experts.
  • Context Length: Boasts an impressive context window of 262,144 tokens.
  • Vision Encoder: Includes a 27-layer vision encoder with Qwen3VL merger, patch 16, and 768 image size, enabling robust multimodal processing.
  • Tokenizer: Features a comprehensive tokenizer with a 248,320 vocabulary and 247,587 BPE merges, along with a chat template.
  • Conversion Details: The text model was dequantized from 733 GGUF Q8_K_P tensors to bf16, and the vision encoder's 334 GGUF F16 tensors were merged.

Ideal Use Cases

  • vLLM Speculation: Specifically prepared for use as a verifier/target model in vLLM speculation setups.
  • Multimodal Applications: Its integrated vision encoder and multimodal architecture make it suitable for tasks requiring both text and image understanding.
  • High-Context Tasks: The extensive context length supports applications demanding processing of very long inputs.