SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked
SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked is a 27.36 billion parameter Qwen3.8 derivative, engineered by SwissNeuron in Switzerland. This BF16 model features a hybrid architecture with 64 language layers and a 1,048,576-token context window achieved via factor-4 YaRN scaling. It is designed for direct technical work and strong reasoning, utilizing capability-preserving post-training and a conservative derisking procedure to maintain instruction fidelity.
Loading preview...
SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked Overview
SwissNeuron is a 27.36 billion parameter Qwen3.8 derivative developed by SwissNeuron in Switzerland. This BF16 model is engineered for direct technical work and strong reasoning, emphasizing precise provenance and reproducible artifacts. It features a hybrid architecture with 64 language layers (48 Gated-DeltaNet and 16 full-attention layers) and an extended 1,048,576-token context window, achieved through factor-4 YaRN scaling over the original 262,144-token window.
Key Capabilities & Differentiators
- Capability-Preserving Post-Training: Utilizes focused internal training and a conservative, low-strength geometric derisking procedure to alter behavior while minimizing impact on reasoning quality and instruction fidelity.
- 1M Context Configuration: Supports a 1,048,576-token context window, though quality at extreme ends may vary as it has not undergone dedicated long-context adaptation.
- Native Thinking Mode: Supports
enable_thinking=Truefor reasoning-intensive tasks andenable_thinking=Falsefor lower-latency direct answers, managed via the Qwen chat template. - Architectural Integrity: Built from the official Qwen3.8-27B base, avoiding merges with aggressively modified checkpoints to preserve core capabilities.
- Multimodal Support: Retains the original MTP (Multimodal Processor) and tokenizer.
When to Use This Model
SwissNeuron is ideal for applications requiring:
- Strong Reasoning: Its design prioritizes maintaining and enhancing reasoning capabilities.
- Technical Work: Engineered for direct technical problem-solving.
- Long Context Processing: Suitable for tasks benefiting from a 1M token context window, with validation recommended for extreme-context use cases.
- Behavioral Control: For users needing specific behavioral alterations without sacrificing core model performance, achieved through its unique derisking process.
Limitations
Users should note that while the 1M context is configured, comprehensive validation for extreme-context quality is ongoing. The model is a full BF16 release, requiring substantial accelerator memory. Outputs should be independently verified for consequential results.