Z.ai's native multimodal agent model for vision-based coding and agent workflows. This is the standard non-thinking variant for image, video, and text inputs, tuned for perceive-plan-execute loops, complex coding, and tool-driven task execution.
Added Apr 1, 2026
Context Window
202.8K
Max Output
131.1K
Input Price (Auto)
$1.20/1M
Output Price (Auto)
$4.00/1M
Cache Read (Auto)
$0.24/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from LMArena.
Arena Score
1433.5
Overall Rank
#93 / 395
Votes
9,373
Confidence Interval
1426.8 - 1440.2
Category Scores
Coding
#78 / 390
2,616 votes
1490.0
Math
#70 / 379
444 votes
1443.2
Longer Query
#88 / 373
4,212 votes
1443.6
Creative Writing
#96 / 393
1,651 votes
1400.4
Instruction Following
#91 / 395
3,249 votes
1423.6
Hard Prompts
#95 / 395
6,190 votes
1453.0
Additional Categories20
English
#73 / 395
3,996 votes
1452.6
Korean
#75 / 265
212 votes
1388.1
Spanish
#75 / 277
290 votes
1435.4
Hard Prompts English
#76 / 393
2,646 votes
1468.4
Chinese
#79 / 367
605 votes
1478.4
Expert
#84 / 345
1,037 votes
1462.6
French
#86 / 276
395 votes
1451.9
Industry Software And It Services
#86 / 395
3,730 votes
1473.1
Industry Writing And Literature And Language
#86 / 394
2,390 votes
1412.5
Polish
#86 / 219
187 votes
1427.4
Industry Mathematical
#87 / 374
520 votes
1440.0
Industry Business And Management And Financial Operations
#90 / 388
1,951 votes
1439.2
Exclude Ties
#93 / 395
6,899 votes
1427.1
Non English
#95 / 395
5,375 votes
1414.6
Industry Entertainment And Sports And Media
#98 / 393
2,197 votes
1396.5
Multi Turn
#98 / 393
1,559 votes
1433.6
Industry Life And Physical And Social Science
#100 / 393
1,464 votes
1448.6
Russian
#101 / 359
950 votes
1421.4
Industry Legal And Government
#123 / 368
732 votes
1423.7
Industry Medicine And Healthcare
#138 / 363
646 votes
1428.6
Published 2026-08-27 · Matched as glm-5v-turbo
LMArena DatasetProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare GLM 5V Turbo with similar models from the same provider or model family.
GLM 5V Turbo Thinking
z-ai/glm-5v-turbo:thinkingThinking-enabled GLM 5V Turbo for image, video, and text inputs. Uses the same multimodal foundation model with more deliberate vision-grounded analysis, planning, and tool use.
GLM 5 Turbo
z-ai/glm-5-turboFast GLM 5 Turbo variant from Z-AI for general chat, coding, and tool use.
GLM 5.3 Flash Uncensored
z-ai/glm-5.3-flash-uncensoredGLM 5.3 Flash Uncensored is an uncensored fine-tune of the efficient 320B mixture-of-experts reasoning model, built for unrestricted chat, creative writing, coding, agentic work, tool use, and long-context tasks.
GLM 5.3 Flash
z-ai/glm-5.3-flashox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.
GLM 4.5V
z-ai/glm-4.5vMultimodal GLM 4.5V that handles images alongside text while keeping the balanced reasoning strength of the GLM 4.5 family.
GLM 4.5V Thinking
z-ai/glm-4.5v:thinkingThinking-enabled GLM 4.5V that surfaces structured reasoning before its final answer. Great for image-grounded analysis, OCR, charts, and deliberate step-by-step responses.
