Ling 3.1 Flash jumps from 3% to 62% on automation benchmarks in one generation
As seen on the 24/7 Wall St. homepage on October 6, 2026.
A jump from 3% to 62% on automation tasks in a single model generation shows how fast low-cost open models are closing in on the agent workloads US labs are charging premium prices for.
Ling 3.1 Flash scores 33% on Terminal-Bench v4.0 and 62% on AutomationBench-AA, up from 0% and 3% for Ling 3.0 Flash. https://t.co/SA9zWZj2Ek
- Replies1
- Reposts0
- Likes5
Continue ReadingShow less
Artificial Analysis posted benchmark results on October 6, 2026, showing that Ling 3.1 Flash scored 62% on AutomationBench-AA.
The predecessor, Ling 3.0 Flash, scored just 3% on AutomationBench-AA. A single model generation erased essentially the entire gap between that baseline and genuinely competitive performance on agentic tasks.
AutomationBench-AA and Terminal-Bench v4.0 are designed to measure the kind of multi-step, tool-using work that AI agents perform in real workflows, the same category of tasks that US frontier labs have been pitching as a premium capability. A low-cost open model closing in on those scores in one release cycle compresses the pricing power that has justified higher subscription tiers.
The speed of this improvement is what investors in AI infrastructure and application companies should be tracking. If capable agent-grade models become broadly available at lower cost, the economics of building on top of them shift quickly, and the competitive moat tied to performance leadership narrows faster than most roadmaps assumed.