Provider logo

GLM 4 32B 0414

THUDM/GLM-4-32B-0414
Provider logo

GLM 4 32B 0414

THUDM/GLM-4-32B-0414

Features 32 billion parameters. Performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series. Pre-trained on 15T of high-quality data, including reasoning-type synthetic data. Enhanced performance in instruction following, engineering code, and function calling.

Context Window

128.0K

Max Output

65.5K

Input Price (Auto)

$0.20/1M

Output Price (Auto)

$0.20/1M

Cache Read (Auto)

$0.10/1M

Benchmarks

Performance metrics and benchmarks

No benchmark data is available yet for this model.

Providers

Auto routing is available for this model. Explicit provider selection is not available.

Loading provider options…

Compare GLM 4 32B 0414 with similar models from the same provider or model family.

GLM 4 9B 0414

THUDM/GLM-4-9B-0414

A 9B parameter version of the GLM-4 series, offering a balance of performance and efficiency.

GLM Z1 9B 0414

THUDM/GLM-Z1-9B-0414

9B small-sized model maintaining the open-source tradition. Despite its smaller scale, GLM-Z1-9B-0414 still exhibits excellent capabilities in mathematical reasoning and general tasks. Its overall performance is already at a leading level among open-source models of the same size.

GLM 5.3 Flash TEE

TEE/glm-5.3-flash

GLM-5.3 Flash is Z.AI's natively multimodal 320B MoE reasoning model with 18B active parameters, served by Phala inside a Trusted Execution Environment with Redpill attestation and signed completion receipts.

GLM 5.3 Flash Uncensored

z-ai/glm-5.3-flash-uncensored

GLM 5.3 Flash Uncensored is an uncensored fine-tune of the efficient 320B mixture-of-experts reasoning model, built for unrestricted chat, creative writing, coding, agentic work, tool use, and long-context tasks.

GLM 5.3 Flash

z-ai/glm-5.3-flash

ox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.

GLM 5.3

zai-org/glm-5.3

GLM-5.3 for long-horizon autonomous coding and engineering workflows. This variant defaults to the model's low reasoning tier for faster responses.