The biggest news on NVIDIA’s conference call tonight is the company revealed they expect to record $20 billion in CPU sales this year. That would set the company up to be the world’s leading CPU supplier.
Yet, that number may actually undersell the opportunity. NVIDIA’s CEO Jensen Huang clarifed later on the call that the $20 billion figure is just for standalone CPU sales. That’s an important distinction because the company sells standalone CPUs in addition to larger systems that include CPUs.
Put succinctly, NVIDIA’s actual CPU revenue will likely be much higher than $20 billion. Here’s the full exchange:
BofA Securities, Research Division
Jensen, there’s a lot of excitement around CPU for Agentic applications and just a lot of noise around the number of CPUs actually exceeding the number of GPUs. And I was just hoping that you could kind of give your perspective that, first of all, is this an incremental workload? Is this kind of cannibalizing what the GPU would have done otherwise?
And then secondly, the $20 billion number that you gave, is that for stand-alone Vera CPUs? Or is that kind of already included in that Vera as part of Vera Rubin? So just if you could educate us on the role of CPU versus GPU is it cannibalistic? Is it incremental? And then the $20 billion number, how to kind of put that in context with what you sell, right, which is usually the CPU as part of the GPU.
Chief Executive Officer
The $20 billion is for stand-alone CPU. And remember, we have Vera is used in 3 ways. As a stand-alone — 4 ways — let me just start with the one that you already know. The first way is Vera Rubin. And we’ll sell millions of Rubins, and every 2 of them is connected to a Vera. And of course, we price those 2 and they’re properly priced. And so that’s #1 use case.
The second use case is Vera standalone CPU. The third is Vera with [ CX-9 ] and the software stack for storage. And then Vera in a [ CX-9 ] — with a software stack for security and compute isolation and confidential computing. Okay, so each one of those use cases is built on Vera. And my sense is that we’ll be supply constrained throughout the entire life of Vera Rubin. There are 4 different use cases of it. And — but anyhow, the answer to your question is of the $20 billion is a stand-alone. With respect to CPUs, an agent is essentially what people call a harness. The agent has a harness that does the — and the harness could be open cloud, it could be Hermes, code — Cloud Code is essentially a harness around Cloud around the OPUS model. OpenAIs Codex is a harness around the GPT 5.5 model.
And so these are harnesses. And these harnesses provide for things like IO, orchestration, memory management tool use connected to tools, for example, browsers and things like that, see compilers, python compilers. And so the harness runs on CPU. And the tool use runs on CPUs. So for example, if the AI were to do a search or do a browser, use a browser that would run on the CPU. The world has 1 billion users, human users. My sense is that the world is going to have billions of agents. Not today, I mean, we’re going to grow into it but we’ll have billions of agents. And those billions of agents will all use tools. And those tools that can be like PCs, just like us humans using PCs today. In the future, you’ll have an agent using PC and so if you kind of think along the lines of in the future, you pick your favorite number of agents at the moment at the moment, call it, a few hundred thousand, but in the future, call it, eventually a few billion.
I could imagine them all using the effectively having PCs that they can all use. And so — but the large length, every one of those agents are going to spin off subagents. And every time they spin these off, you’re going to need to do inference. That’s where the thinking happens. All of the thinking happens on GPUs, all of the orchestration essentially runs on CPUs. And the subagents when they’re spun off, they — when they’re thinking they use GPUs. Whenever the agents use simulators, those can run on CPUs or GPUs, which is the reason why we’re working so closely with Cadence and Synopsys and to accelerate all of the world’s tool we’re accelerating all of the world’s tools and data processing engines and database engines because agents use these tools and have — they have lower patients tolerance humans, and they want things to happen quickly. And so we’re accelerating all the world’s tools so that it runs on CUDA. And you could see us doing that when I work with Cadence and Synopsys, and Siemens and companies in Adobe. And that’s because we’re trying to get all of the world’s tools to run on GPUs because they already have GPUs, and it’s a lot faster. So we’re going to need a lot more CPUs, and Vera was designed to be an agenetic CPU. The CPUs of the past were designed to have many cores so that it could be easily rentable. People rent at course. Well, agents don’t rent cores.
They just want the work to be done fast. The economics of the past was dollars per core. That’s the economics of cloud computing of the past. The economics of the AI of the future is tokens per dollar or dollars per token. And so what we need to do in the future is to generate tokens, process tokens as fast as possible, and that’s what Vera does incredibly well. So we’re expecting to be very successful with Vera. But ultimately, ultimately, what we’re doing is we’re building infrastructure for AI and it needs incredibly great storage.
That’s the reason why we built STX. It needs incredibly good networking. That’s why we have Spectrum-X. It needs incredibly great GPUs, of course, and inferencing ability. That’s the reason why NVLink 72. It needs incredibly great security and confidential computing, which is the reason why Vera Rubin is the world’s first platform with end-to-end confidential computing and it needs great CPUs. We’ve got it all covered.