zai-org/GLM-4.7

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:4Model Size:357BQuant:FP8Context Size:198kTool Calling:SupportedPublished:Dec 22, 2025License:mitArchitecture:Transformer2.0K Open Weights Warm

GLM-4.7 is a powerful language model developed by zai-org, specifically engineered for advanced coding tasks and complex reasoning. It demonstrates significant improvements in multilingual agentic coding, terminal-based operations, and mathematical problem-solving, achieving 73.8% on SWE-bench and 42.8% on HLE (with tools). The model also features enhanced tool-using capabilities and innovative thinking modes like Interleaved and Preserved Thinking, making it highly effective for intricate, multi-turn agentic workflows.

Loading preview...

GLM-4.7: Advanced Coding and Reasoning

GLM-4.7, developed by zai-org, is an instruction-tuned language model designed to excel in coding, reasoning, and tool-using tasks. It represents a substantial upgrade over its predecessor, GLM-4.6, particularly in agentic coding and complex problem-solving.

Key Capabilities & Features

  • Core Coding: Achieves 73.8% on SWE-bench and 66.7% on SWE-bench Multilingual, with significant gains in terminal-based tasks (41% on Terminal Bench 2.0). It supports thinking before acting in agent frameworks like Claude Code and Kilo Code.
  • Vibe Coding: Improves UI quality, generating cleaner webpages and better-looking slides with accurate layouts.
  • Tool Using: Shows marked improvements on benchmarks like τ²-Bench (87.4%) and web browsing via BrowseComp (52.0%).
  • Complex Reasoning: Delivers a substantial boost in mathematical and reasoning capabilities, scoring 42.8% on the HLE (Humanity’s Last Exam) benchmark with tools.
  • Innovative Thinking Modes: Enhances Interleaved Thinking (thinking before every response/tool call) and introduces Preserved Thinking (retains thinking blocks across multi-turn conversations for coding agents) and Turn-level Thinking (per-turn control over reasoning).

Performance Highlights

GLM-4.7 demonstrates competitive performance across various benchmarks, including MMLU-Pro (84.3%), GPQA-Diamond (85.7%), and AIME 2025 (95.7%), often outperforming or matching models like DeepSeek-V3.2 and Claude Sonnet 4.5 in specific coding and reasoning metrics.

Ideal Use Cases

  • Software Development: Excellent for code generation, debugging, and agentic coding workflows, especially in multilingual environments.
  • Complex Problem Solving: Suited for tasks requiring advanced mathematical and logical reasoning.
  • Automated Agents: Its enhanced tool-using and thinking modes make it ideal for building robust, multi-turn AI agents that require consistent reasoning.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p