Ling-3.0-flash Thinking enables visible reasoning on inclusionAI's token-efficient 124B-parameter Mixture-of-Experts model for harder coding, tool use, planning, and production-scale agent workflows.
Added Jul 23, 2026
Model weightsContext Window
262.1K
Max Output
32.8K
Input Price (Auto)
$0.075/1M
Output Price (Auto)
$0.22/1M
Cache Read (Auto)
$0.015/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
37.8
Coding Index
50.6
Agentic Index
29.3
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
85.5%
Better than 82% of models compared
HLE
Humanity's Last Exam
23.7%
Better than 79% of models compared
AA-LCR
Long context reasoning evaluation
67.0%
Better than 73% of models compared
GDPval-AA
Economically valuable tasks
30.3%
CritPt
Research-level physics reasoning
1.7%
Coding
SciCode
Python programming for scientific computing
41.1%
Better than 75% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
18.2%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
44.1%
Last updated Aug 18, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Ling 3.0 Flash Thinking with similar models from the same provider or model family.
Ling 3.0 Flash
inclusionai/ling-3.0-flashLing-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.
GLM 5.3 Flash TEE
TEE/glm-5.3-flashGLM-5.3 Flash is Z.AI's natively multimodal 320B MoE reasoning model with 18B active parameters, served by Phala inside a Trusted Execution Environment with Redpill attestation and signed completion receipts.
GLM 5.3 Flash Uncensored
z-ai/glm-5.3-flash-uncensoredGLM 5.3 Flash Uncensored is an uncensored fine-tune of the efficient 320B mixture-of-experts reasoning model, built for unrestricted chat, creative writing, coding, agentic work, tool use, and long-context tasks.
Qwen3.8 Flash
alibaba/qwen3.8-flashQwen3.8 Flash is Alibaba's latest fast multimodal model, with a million-token context window for coding, agentic workflows, visual understanding, long documents, codebases, and videos.
GLM 5.3 Flash
z-ai/glm-5.3-flashox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.
DeepSeek V4 Flash Vision Exp
deepseek/deepseek-v4-flash-vision-expAn experimental vision-enabled DeepSeek V4 Flash model that adds image understanding while retaining the text, reasoning, coding, tool-calling, and agent capabilities of the base model. This route is served directly by DeepSeek, so privacy and logging guarantees are limited.
