Rupok/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 12, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Rupok/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled is a 27 billion parameter language model fine-tuned on the Qwen3.5 architecture, leveraging Chain-of-Thought (CoT) distillation from Claude-4.6 Opus interactions. It excels at structured reasoning, breaking down complex problems, and planning step-by-step solutions within tags. This model is optimized for analytical tasks, coding, and mathematics, providing improved autonomy and stability in agent environments.

Loading preview...

Model Overview

Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled is a 27 billion parameter model built upon the Qwen3.5 architecture, specifically fine-tuned for advanced reasoning capabilities. It incorporates state-of-the-art Chain-of-Thought (CoT) distillation, primarily from Claude-4.6 Opus interactions, to enhance its problem-solving methodology.

Key Capabilities

  • Structured Reasoning: The model is designed to break down complex problems, plan solutions step-by-step within <think> tags, and deliver precise answers. It adopts an efficient reasoning paradigm, reducing redundant cognitive loops.
  • Agent Environment Optimization: It offers native support for the "developer" role and preserves its thinking mode, allowing for continuous autonomous operation for over 9 minutes in coding agent environments like Claude Code and OpenCode. This significantly improves autonomy and stability compared to the base model.
  • Efficient Fine-tuning: Developed using Unsloth, the model's fine-tuning focused on injecting high-density reasoning logic and enforcing a strict output format, with loss calculated purely over <think> sequences and solutions.
  • Hardware Efficiency: Achieves 29–35 tokens/second generation speed with approximately 16.5 GB VRAM using Q4_K_M quantization, supporting a full 262K context.

Ideal Use Cases

  • Offline Analytical Tasks: Best suited for scenarios requiring deep analysis and transparent internal logic.
  • Coding and Mathematics: Excels in coding-agent environments and tasks demanding heavy logic-dependent prompting.
  • Structured Problem Solving: Useful for applications where step-by-step reasoning and planning are critical.