glm-5.2

glm-5.2

Zhipu AI’s next-generation, self-developed, high-end flagship large model specifically designed for multimodal agents.
2026-06-15
LLM
Model capability: thinkingModel capability: function_call
Cache Read:
$0.26/1M tokens
Input:
$1.4/1M tokens
Output:
$4.4/1M tokens
Bulk order? Contact your manager for exclusive deals

API Overview

The Glm-5.2 model is a next-generation, self-developed, high-end flagship large-scale model created by Zhipu AI specifically for multimodal intelligent agents. The model sets new industry benchmarks in key areas such as long-range academic reasoning, code engineering, and collaborative orchestration of tens of thousands of APIs. As a culmination of native multimodal capabilities, Glm-5.2 not only demonstrates cutting-edge performance on multilingual complex texts but also exhibits exceptional autonomous evolution capabilities—both on-device and in the cloud—in the fields of “end-to-end streaming fusion interaction among audio, vision, and text” and “ultra-realistic emotion and dialect multi-track control.” It is designed specifically for the next-generation enterprise-level full-scenario deployment and ultra-realistic digital interaction. ───────────────────────────────────────────────────────────────────

Core Capabilities

Ultra-realistic end-to-end multimodal fusion: Adopts an all-modal integrated architecture. Supports end-to-end bidirectional streaming interaction involving text, high-precision images, real-time video streams, and ultra-realistic speech. In multimodal real-time training and audio-video collaborative diagnosis scenarios, it significantly reduces the first-word latency (TTFT) and boasts exceptional abilities in capturing emotional nuances, controlling breathing sounds, and understanding regional dialectal intonations and metaphors. Million-scale tool manipulation and long-term scenario planning: Tailored for complex enterprise workflows, the model has been deeply enhanced for tool invocation and autonomous orchestration of long-term tasks. Within a single round of complex agent workflows, it can precisely retrieve, schedule, and闭环 execute hundreds or even thousands of enterprise-grade private API interfaces, perfectly handling highly complex automated financial compliance audits and cross-platform supply chain management. Million-token-scale retrieval of massive long texts: By default, it supports an enormous context window of up to 1 million tokens, maintaining a 100% lossless information recall rate in full-scale complex document “needle-in-a-haystack” tests. It can effortlessly process vast corporate financial reports, technical codebases, or massive historical archives spanning multiple directories, containing millions of words, and perform extremely precise logical drill-downs and compliance audits. Production-level high concurrency and highly optimized computational operators: Thanks to the newly upgraded mixed-expert (MoE) routing algorithm and self-developed high-performance computing operators, the model delivers flagship-level intelligence while significantly boosting overall throughput. It demonstrates outstanding resilience under high-frequency, complex business workloads, ensuring high output stability and ultra-fast response speeds to support large-scale enterprise readiness.

Playground

Log in to explore more features! Click to Log In

API Analytics

API Reference (1)

API DescriptionAPI EndpointRequest MethodStabilityParameter Description
Chat (Zhipu GLM Multimodal)
POST
Stable
View Details

API Pricing

$
ModelDescriptionContextOfficial Price302.AI PriceCache Price

glm-5.2

-
1000000

Input$1.4 / 1M tokens
Output$4.4 / 1M tokens

Input$1.4/ 1M tokens
Output$4.4/ 1M tokens
Original Price

Cache Read$0.26/ 1M tokens