ox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.
Added Aug 26, 2026
Model weightsContext Window
1.0M
Max Output
131.1K
Input Price (Auto)
$0.075/1M
Output Price (Auto)
$0.25/1M
Cache Read (Auto)
$0.015/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
57.5
Coding Index
71.5
Agentic Index
58.2
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
91.2%
Better than 95% of models compared
HLE
Humanity's Last Exam
39.9%
Better than 93% of models compared
AA-LCR
Long context reasoning evaluation
78.0%
Better than 96% of models compared
GDPval-AA
Economically valuable tasks
63.2%
CritPt
Research-level physics reasoning
15.4%
Coding
SciCode
Python programming for scientific computing
46.1%
Better than 85% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
27.5%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
27.6%
Last updated Aug 28, 2026, 12:00 PM
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare GLM 5.3 Flash with similar models from the same provider or model family.
GLM 5.3 Flash Uncensored
z-ai/glm-5.3-flash-uncensoredGLM 5.3 Flash Uncensored is an uncensored fine-tune of the efficient 320B mixture-of-experts reasoning model, built for unrestricted chat, creative writing, coding, agentic work, tool use, and long-context tasks.
GLM 5V Turbo Thinking
z-ai/glm-5v-turbo:thinkingThinking-enabled GLM 5V Turbo for image, video, and text inputs. Uses the same multimodal foundation model with more deliberate vision-grounded analysis, planning, and tool use.
GLM 5V Turbo
z-ai/glm-5v-turboZ.ai's native multimodal agent model for vision-based coding and agent workflows. This is the standard non-thinking variant for image, video, and text inputs, tuned for perceive-plan-execute loops, complex coding, and tool-driven task execution.
GLM 5 Turbo
z-ai/glm-5-turboFast GLM 5 Turbo variant from Z-AI for general chat, coding, and tool use.
GLM 4.5V
z-ai/glm-4.5vMultimodal GLM 4.5V that handles images alongside text while keeping the balanced reasoning strength of the GLM 4.5 family.
GLM 4.5V Thinking
z-ai/glm-4.5v:thinkingThinking-enabled GLM 4.5V that surfaces structured reasoning before its final answer. Great for image-grounded analysis, OCR, charts, and deliberate step-by-step responses.
