Provider logo

Ling 3.0 Flash

inclusionai/ling-3.0-flash
Provider logo

Ling 3.0 Flash

inclusionai/ling-3.0-flash

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.

Added Jul 23, 2026

Model weights

Context Window

262.1K

Max Output

32.8K

Input Price (Auto)

$0.075/1M

Output Price (Auto)

$0.22/1M

Cache Read (Auto)

$0.015/1M

Capabilities

Benchmarks

Performance metrics and benchmarks

Sourced from Artificial Analysis.

Intelligence Index

37.8

Better than 81% of models compared

Coding Index

50.6

Better than 65% of models compared

Agentic Index

29.3

Better than 65% of models compared

Reasoning

GPQA Diamond

Graduate-level scientific reasoning

85.5%

Better than 82% of models compared

HLE

Humanity's Last Exam

23.7%

Better than 79% of models compared

AA-LCR

Long context reasoning evaluation

67.0%

Better than 73% of models compared

GDPval-AA

Economically valuable tasks

30.3%

CritPt

Research-level physics reasoning

1.7%

Coding

SciCode

Python programming for scientific computing

41.1%

Better than 75% of models compared

Knowledge

AA-Omniscience Accuracy

Proportion of correctly answered questions

18.2%

AA-Omniscience Hallucination Rate

Rate of incorrect answers among non-correct responses

44.1%

Last updated Aug 18, 2026

Artificial Analysis

Providers

Auto routing is available for this model. Explicit provider selection is not available.

Loading provider options…

Compare Ling 3.0 Flash with similar models from the same provider or model family.

Ling 3.0 Flash Thinking

inclusionai/ling-3.0-flash:thinking

Ling-3.0-flash Thinking enables visible reasoning on inclusionAI's token-efficient 124B-parameter Mixture-of-Experts model for harder coding, tool use, planning, and production-scale agent workflows.

GLM 5.3 Flash TEE

TEE/glm-5.3-flash

GLM-5.3 Flash is Z.AI's natively multimodal 320B MoE reasoning model with 18B active parameters, served by Phala inside a Trusted Execution Environment with Redpill attestation and signed completion receipts.

GLM 5.3 Flash Uncensored

z-ai/glm-5.3-flash-uncensored

GLM 5.3 Flash Uncensored is an uncensored fine-tune of the efficient 320B mixture-of-experts reasoning model, built for unrestricted chat, creative writing, coding, agentic work, tool use, and long-context tasks.

Qwen3.8 Flash

alibaba/qwen3.8-flash

Qwen3.8 Flash is Alibaba's latest fast multimodal model, with a million-token context window for coding, agentic workflows, visual understanding, long documents, codebases, and videos.

GLM 5.3 Flash

z-ai/glm-5.3-flash

ox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.

DeepSeek V4 Flash Vision Exp

deepseek/deepseek-v4-flash-vision-exp

An experimental vision-enabled DeepSeek V4 Flash model that adds image understanding while retaining the text, reasoning, coding, tool-calling, and agent capabilities of the base model. This route is served directly by DeepSeek, so privacy and logging guarantees are limited.