zai-org/GLM-5.3-Flash

Hugging Face
TEXT GENERATIONPricing:Input $0.15 / Cached $0.03 / Output $0.5Concurrent Unit Cost:4Model Size:321.3BQuant:FP8Context Size:256kTool Calling:SupportedPublished:Aug 25, 2026License:mitArchitecture:Transformer1.3K Open Weights Warm

GLM-5.3-Flash is a 320 billion parameter natively multimodal model from zai-org, featuring 18 billion active parameters and a 32768 token context length. It utilizes a hybrid architecture combining sparse and linear attention, alongside Manifold-Constrained Hyper-Connections (mHC), to reduce long-context serving costs while maintaining capabilities. This model is designed for high intelligence and efficiency, outperforming previous GLM versions and approaching top-tier models on coding and agentic benchmarks.

Loading preview...

GLM-5.3-Flash Overview

GLM-5.3-Flash is the latest iteration in the GLM-5 series by zai-org, distinguished as its first natively multimodal model. With a total of 320 billion parameters, it operates efficiently with only 18 billion active parameters, offering significant cost reductions compared to prior versions. The model's architecture has been re-engineered for enhanced capability and efficiency, incorporating a novel hybrid approach that blends sparse and linear attention to optimize long-context processing costs without sacrificing performance. It also integrates Manifold-Constrained Hyper-Connections (mHC) for improved scaling efficiency.

Key Capabilities & Features

  • Multimodal: Natively supports multimodal inputs, trained on a 30T-token multimodal pre-training corpus.
  • Efficiency: Achieves high performance with only 18B active parameters, leading to a tenfold price reduction compared to GLM-5.2.
  • Advanced Architecture: Features a hybrid sparse and linear attention mechanism for cost-effective long-context handling and Manifold-Constrained Hyper-Connections (mHC) for scaling.
  • Strong Performance: Demonstrates superior performance over GLM-5.2 across various benchmarks and real-world tasks, closely matching models like Claude Opus 4.8 on coding and agentic evaluations.
  • Configurable Reasoning: Supports a reasoning_effort parameter with low, high, and max settings to control computational budget.
  • Local Deployment: Compatible with popular frameworks such as SGLang, vLLM, TokenSpeed, Transformers, KTransformers, and Unsloth.

Use Cases & Strengths

GLM-5.3-Flash is particularly well-suited for applications requiring advanced coding, agentic reasoning, and multimodal understanding. Its efficiency and strong benchmark performance make it a compelling choice for developers seeking powerful yet cost-effective large language models for complex tasks.