Ling 3.1 Flash jumps from 3% to 62% on automation benchmarks in one generation

As seen on the 24/7 Wall St. homepage on October 6, 2026.

Continue ReadingShow less

Artificial Analysis posted benchmark results on October 6, 2026, showing that Ling 3.1 Flash scored 62% on AutomationBench-AA.

The predecessor, Ling 3.0 Flash, scored just 3% on AutomationBench-AA. A single model generation erased essentially the entire gap between that baseline and genuinely competitive performance on agentic tasks.

AutomationBench-AA and Terminal-Bench v4.0 are designed to measure the kind of multi-step, tool-using work that AI agents perform in real workflows, the same category of tasks that US frontier labs have been pitching as a premium capability. A low-cost open model closing in on those scores in one release cycle compresses the pricing power that has justified higher subscription tiers.

The speed of this improvement is what investors in AI infrastructure and application companies should be tracking. If capable agent-grade models become broadly available at lower cost, the economics of building on top of them shift quickly, and the competitive moat tied to performance leadership narrows faster than most roadmaps assumed.