For three years the AI infrastructure conversation has been almost entirely about GPUs. That framing is now out of date, and the delivery of NVIDIA's first custom CPU to the world's top AI labs this week is a useful moment to understand why. AI agents are not the same as chatbots. A chatbot takes a question and generates an answer. An agent takes a goal, breaks it into steps, calls tools, queries databases, writes and runs code, checks the output, and loops back. The GPU handles the actual model inference, but everything between those inference steps, the orchestration, the tool calls, the data retrieval, the policy checks, falls on the CPU. Research from Georgia Tech and Intel put a number on this: CPU-side processing accounts for between 50% and 90% of total latency in agentic workloads. In many agent pipelines, the GPU is simply waiting for the CPU to catch up. NVIDIA's answer is the Vera CPU, now in production and physically delivered to Anthropic, OpenAI, SpaceXAI, and Oracle Cloud this week. The chip carries 88 custom cores, 1.2 TB/s of memory bandwidth, and is built on Arm architecture rather than the x86 designs that Intel and AMD have long dominated. NVIDIA claims it runs 50% faster than traditional server CPUs under full load, at twice the energy efficiency. Oracle has already committed to deploying hundreds of thousands of Vera CPUs in 2026, the largest cloud commitment to the chip so far. That number signals the scale of what is coming. Oracle is betting that enterprise customers will want production-grade agentic infrastructure that they cannot get elsewhere today. NVIDIA is not alone in sensing this shift. AMD's CEO called surging CPU demand a major growth driver on her last earnings call, with the company raising its server CPU market estimate from $60 billion to over $120 billion by 2030, doubling its own forecast in under six months. Intel reportedly has a six-month order backlog. Arm launched its own data center CPU chip on March 25th, just days after NVIDIA's GTC announcement, breaking a 35-year history of selling only chip designs rather than finished chips. The structural change in CPU-to-GPU ratios is what makes this worth watching closely. Traditional AI clusters ran roughly one CPU for every four to eight GPUs. Agentic workloads are pushing that ratio toward one-to-one or even CPU-heavy configurations. Arm estimates that CPU core demand per gigawatt of data center power will quadruple as agentic AI scales. For anyone running or procuring cloud services, this has a direct cost implication. Cloud providers will need to build significantly more CPU capacity to support agentic workloads, and those costs will flow through to customers. The businesses deploying agents at scale, whether for customer service, back-office automation, or software development, will find that agent speed and cost are increasingly determined by CPU infrastructure, not just by which AI model they chose. NVIDIA's entry into the CPU market also has competitive implications that go beyond AI. By building its own CPU, it reduces its dependence on Intel and AMD to build the host processors that sit alongside its GPUs. Vera plugs directly into NVIDIA's own GPU racks via a high-bandwidth connection that traditional CPUs cannot match. The entire system becomes NVIDIA's to sell, price, and optimize. The server CPU market was worth $26 billion in 2025. If Morgan Stanley's projections are right, it could reach $60 billion to $100 billion by 2030. The companies positioned to capture that growth are NVIDIA, AMD, Arm, and the cloud providers building their own custom chips. Intel is fighting to stay relevant. The businesses that depend on AI agents to operate, and that number is growing fast, will ultimately pay for whatever infrastructure wins.
Infrastructure2 min read
NVIDIA Ships Its First Custom CPU to OpenAI, Anthropic, Oracle
June 5, 2026Synthesized from 2 sources: AI News, NVIDIA
NVIDIA has delivered the first units of its Vera CPU, a custom chip built specifically for AI agents, to Anthropic, OpenAI, SpaceXAI, and Oracle, marking the moment that a quiet but significant shift in AI infrastructure, from GPU-dominated to CPU-heavy, moves from theory to production hardware.
Related Coverage
Stay informed
Get AI intelligence like this delivered to your inbox.