DeepSeek V4.1 Flash beats the 1.6T Pro model at a quarter of the cost

As seen on the 24/7 Wall St. homepage on September 10, 2026.

A model roughly a third the size beating DeepSeek's own 1.6T flagship at about a quarter of the cost per token undercuts the compute-scaling logic funding the AI infrastructure trade.

DeepSeek V4.1 Flash overtakes DeepSeek V4 Pro 0813 as DeepSeek’s new flagship model with a score of 40 on Artificial Analysis Intelligence Index. At just 552B parameters, it outperforms the Pro (1.6T) model while costing ~4x less per token, placing it just short of the https://t.co/ne5qNtijJ5
  • Replies12
  • Reposts11
  • Likes114
Continue ReadingShow less

DeepSeek V4.1 Flash has displaced DeepSeek V4 Pro 0813 as the company's top-ranked model on the Artificial Analysis Intelligence Index, doing it with 552 billion parameters against the Pro's 1.6 trillion.

V4.1 Flash runs at roughly a quarter of the price per token compared to the model it just dethroned. Getting superior benchmark performance at that kind of discount removes one of the core arguments for deploying the heavier, more expensive architecture.

Disclosure

*$149 for two years (or $1.43 per week) is an introductory promotion for new members only. 62% discount based on the current list price of Stock Advisor of $199/year. Membership will renew at the then-current list price at the end of the membership term. Stock Advisor returns are 930% as compared to the S&P 500 returns of 185% as of April 7, 2026.

The same investor newsletter that told subscribers to buy Amazon in 2002, Netflix in 2004, and Nvidia in 2005 still publishes two new stock picks every month. Over 23 years, Motley Fool's Stock Advisor has more than quadrupled the S&P 500. New members get this month's picks, the Top 10 Rankings, and a 30-day money-back guarantee. Click here to unlock their next top stocks while new members are still being accepted.

For investors who have been buying into AI infrastructure on the logic that bigger models inevitably require more compute, this result complicates that narrative. A model roughly a third the size outperforming the flagship challenges the assumption that parameter count and capital expenditure scale together reliably.

The tweet from Artificial Analysis trails off before naming what V4.1 Flash falls just short of, so the competitive ceiling here remains unclear from the available data. Efficiency gains at this magnitude tend to ripple across pricing benchmarks industry-wide, putting pressure on providers still charging premium rates for scale.