apurvaga/Qwen3.6-35B-A3B-preserve-thinking

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 12, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

apurvaga/Qwen3.6-35B-A3B-preserve-thinking is a 35.1 billion parameter causal language model based on the Qwen architecture. This model is a specialized version of Qwen/Qwen3.6-35B-A3B, featuring a customized tokenizer configuration and chat template. Its primary differentiator is the preservation of prior assistant reasoning content across conversational turns, making it suitable for use cases requiring consistent, multi-turn reasoning.

Loading preview...

Model Overview

apurvaga/Qwen3.6-35B-A3B-preserve-thinking is a 35.1 billion parameter language model derived from the Qwen/Qwen3.6-35B-A3B base model. While the core model weights remain unchanged, this version incorporates a custom tokenizer configuration and chat template.

Key Differentiator

The most significant feature of this model is its unique chat template, which is designed to preserve prior assistant reasoning content across conversational turns. Unlike standard templates that might discard intermediate reasoning, this model maintains that context, which can be crucial for complex, multi-step interactions.

Intended Use Cases

This model is particularly well-suited for applications where:

  • Multi-turn reasoning is critical: Scenarios requiring the model to build upon its previous thought processes or explanations.
  • Consistent contextual understanding: Maintaining a deeper, more persistent understanding of the conversation's logical flow.
  • Complex problem-solving: Tasks where the model's internal 'thinking' process needs to be carried forward to subsequent responses.