Hastagaras/Qwen3.5-9B-GLM-Wannabe

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 4, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Hastagaras/Qwen3.5-9B-GLM-Wannabe is an experimental 9 billion parameter full finetune of the Qwen3.5-9B model, developed by Hastagaras. It utilizes distilled data from GLM 4.7 and GLM 5, specifically designed to test a pipeline for linear attention training with a gated delta mechanism. This model focuses on exploring advanced attention architectures rather than achieving specific benchmark goals, and was trained with a context length of 16,384 tokens.

Loading preview...

Overview

Hastagaras/Qwen3.5-9B-GLM-Wannabe is an experimental 9 billion parameter model based on Qwen3.5-9B. Its primary purpose is to serve as a testbed for a novel pipeline involving linear attention training, specifically incorporating a gated delta mechanism. The model was finetuned using a unique blend of distilled data from GLM 4.7 and GLM 5, alongside self-generated question-answer pairs.

Key Characteristics

  • Base Model: Qwen3.5-9B, inheriting its native support for up to 262,144 tokens, though training was conducted at 16,384 tokens.
  • Training Data: A mix of multi-turn CoT distillation from Jackrong/glm-4.7-multiturn-CoT, single-turn GLM-5 distilled data from TeichAI/Pony-Alpha-15k, and simple self-generated Q&A pairs.
  • Attention Mechanism: Implements and tests a gated delta mechanism for linear attention, compiled with torch_xla.experimental.scan, alongside standard attention accelerated with torch_xla.experimental.custom_kernel.flash_attention.
  • Preprocessing: Data was reformatted from Gemini/GLM markdown to align with Qwen conventions for improved convergence.

Intended Use

This model is an experimental finetune focused on research and development of attention mechanisms. It has no specific benchmark goals and has not been evaluated on downstream tasks. Developers interested in exploring advanced linear attention architectures and their training pipelines may find this model particularly relevant.