SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked

VISIONPricing:Input $1.06 / Cached $0.15 / Output $2.6Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 23, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked is a 27.36 billion parameter Qwen3.8 derivative, engineered by SwissNeuron in Switzerland. This BF16 model features a hybrid architecture with 64 language layers and a 1,048,576-token context window achieved via factor-4 YaRN scaling. It is designed for direct technical work and strong reasoning, utilizing capability-preserving post-training and a conservative derisking procedure to maintain instruction fidelity.

Loading preview...

SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked Overview

SwissNeuron is a 27.36 billion parameter Qwen3.8 derivative developed by SwissNeuron in Switzerland. This BF16 model is engineered for direct technical work and strong reasoning, emphasizing precise provenance and reproducible artifacts. It features a hybrid architecture with 64 language layers (48 Gated-DeltaNet and 16 full-attention layers) and an extended 1,048,576-token context window, achieved through factor-4 YaRN scaling over the original 262,144-token window.

Key Capabilities & Differentiators

  • Capability-Preserving Post-Training: Utilizes focused internal training and a conservative, low-strength geometric derisking procedure to alter behavior while minimizing impact on reasoning quality and instruction fidelity.
  • 1M Context Configuration: Supports a 1,048,576-token context window, though quality at extreme ends may vary as it has not undergone dedicated long-context adaptation.
  • Native Thinking Mode: Supports enable_thinking=True for reasoning-intensive tasks and enable_thinking=False for lower-latency direct answers, managed via the Qwen chat template.
  • Architectural Integrity: Built from the official Qwen3.8-27B base, avoiding merges with aggressively modified checkpoints to preserve core capabilities.
  • Multimodal Support: Retains the original MTP (Multimodal Processor) and tokenizer.

When to Use This Model

SwissNeuron is ideal for applications requiring:

  • Strong Reasoning: Its design prioritizes maintaining and enhancing reasoning capabilities.
  • Technical Work: Engineered for direct technical problem-solving.
  • Long Context Processing: Suitable for tasks benefiting from a 1M token context window, with validation recommended for extreme-context use cases.
  • Behavioral Control: For users needing specific behavioral alterations without sacrificing core model performance, achieved through its unique derisking process.

Limitations

Users should note that while the 1M context is configured, comprehensive validation for extreme-context quality is ongoing. The model is a full BF16 release, requiring substantial accelerator memory. Outputs should be independently verified for consequential results.