r3lax/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 20, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

r3lax/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled is a 35.1 billion parameter Mixture-of-Experts (MoE) model, developed by r3lax, that distills the reasoning capabilities of Anthropic's Claude Opus 4.7 into a permissively licensed architecture. It is fine-tuned to imitate Claude's chain-of-thought style, including explicit blocks, and operates with an active parameter count of approximately 3 billion per token, offering the capacity of a larger model at lower inference cost. This model excels at complex reasoning tasks such as graduate-level STEM problems, competition mathematics, and multi-step logic puzzles, supporting a 64k token context length for extensive internal thought processes.

Loading preview...

Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled Overview

This model is a reasoning-distilled variant of the Qwen3.6-35B-A3B Mixture-of-Experts (MoE) base, fine-tuned by r3lax to emulate the advanced chain-of-thought reasoning style of Anthropic's Claude Opus 4.7. The primary goal is to bring Claude-grade reasoning behavior into an open-weights, efficiently runnable MoE model. It achieves this by training on approximately 8,000 high-quality reasoning traces from Opus 4.7, teaching the model to explicitly "think" using <think>…</think> blocks before providing a final answer.

Key Capabilities

  • Claude-style Reasoning: Emulates the detailed, step-by-step reasoning process of Claude Opus 4.7, including explicit internal thought processes.
  • Efficient Inference: As a 35.1B parameter MoE with 256 experts (8 routed + 1 shared), it activates only about 3B parameters per token, providing the capacity of a larger model with the inference cost of a smaller dense model.
  • Extended Context for Reasoning: Supports a 64k token context length, allowing for extensive internal reasoning (5-30k tokens of <think> output) on challenging problems.
  • Strong Reasoning Performance: Achieves 84.3% on GSM8K CoT (8-shot) and 74.9% on MMLU-Pro (5-shot), demonstrating proficiency in complex problem-solving.
  • Permissively Licensed: Built on the Apache-2.0 licensed Qwen3.6 base, enabling broad use and further development.

Good For

  • Hard Reasoning Tasks: Excels in domains like graduate-level STEM, competition mathematics (e.g., AIME), code reasoning with explicit walk-throughs, and multi-step logic puzzles.
  • Agentic Planning: Ideal for scenarios where explicit <think> blocks can significantly improve correctness and transparency of decision-making.
  • Research and Development: Provides a clean base and a separately published LoRA adapter for further fine-tuning or stacking with other checkpoints.
  • Resource-Efficient Deployment: Can run full-quality bf16 inference on a single 80GB A100/H100 GPU, and quantized GGUF versions are available for consumer hardware (e.g., IQ4_XS fits in ~24GB RAM/VRAM).