Nvidia’s Vera pairs 88 custom Arm cores with 1.8 TB/s NVLink, aiming at non-GPU AI agents
A standalone CPU built to run AI agent hosts and orchestrate Vera Rubin systems, not just “more compute.”

Nvidia says its Vera CPU, and the Olympus custom Armv9.2 core inside it, is designed for two jobs: managing GPUs in Vera Rubin systems and hosting AI agents. For decision-makers, the bet is simple but disruptive: hyperscalers and cloud providers can buy Nvidia’s CPU stack without buying into its GPUs first.
Nvidia’s Vera is not trying to be “a better CPU.” It is trying to be a necessary CPU. And the most concrete tell is hiding in the networking math: Nvidia’s Vera CPU Superchip connects two Vera CPUs over NVLink-C2C with 1.8 TB/s of bidirectional bandwidth, for a total of 176 cores and 352 threads.
That bandwidth matters because it sets up Vera’s real target workloads. Nvidia’s Vera is positioned for (1) an AI head node role that manages GPUs in its upcoming Vera Rubin systems, and (2) a more contentious role, hosting AI agents. The key nuance is right there: unlike the large language models (LLMs) that power many agents, these AI agent hosts do not actually run on GPUs. So the CPU is doing different work than “feeding the accelerator.” It is coordinating, executing, and keeping pipelines moving while the agent logic churns.
Under the hood, Nvidia’s design choices lean into that coordination problem. Vera is built as a superchip-style package: the monolithic compute die is where the 88 custom Armv9.2-compatible cores live, and Nvidia says this compute die is fabbed on TSMC’s 3nm process. Nvidia argues monolithic compute improves core-to-core bandwidth and latency compared to competing multi-die approaches. Around that compute die, Vera uses chiplets for I/O and memory. The system includes eight LPDDR5x controllers, and the whitepaper describes two distinct I/O dies, one handling PCIe 6.4 and CXL 3.1 connectivity, and another dedicated to the chip’s NVLink Chip-to-Chip interface.
If this sounds a little like Amazon’s Graviton 4 strategy, that is not accidental. The source points out that Vera is “somewhat reminiscent” of Graviton 4: monolithic compute paired with disaggregated I/O and memory functionality. That pattern matters for buyers because it is a roadmap for system design. Memory bandwidth and I/O bandwidth do not behave like optional upgrades in large agent stacks; they are first-order constraints that determine whether expensive accelerators spend their time computing or waiting.
Now zoom in on the Vera CPU scale-up story. Like most modern datacenter CPUs, Vera supports both single- and dual-socket configurations. In dual-socket, Nvidia calls it the Vera CPU Superchip. The source says the Superchip features two Vera CPUs connected over NVLink-C2C at 1.8 TB/s bidirectional bandwidth. It also states the memory configuration: 16 SOCAMM2 LPDDR5x memory modules delivering 2.4 TB/s aggregate memory bandwidth, 1.2 TB/s per side. Nvidia’s agentic AI reference designs then escalate the stack into rack-level planning. The designs call for cramming as many as 128 superchips, or 256 CPUs, totaling 22,528 cores and 384 TB of memory into a single liquid-cooled rack.
That rack math is one reason Vera is positioned as a standalone platform, independent of Nvidia’s GPUs. Nvidia’s launch framing is explicit: for the first time it has directly challenged Intel and AMD’s CPU dominance, aiming to sell its standalone CPUs to hyperscalers and other cloud providers. The source lists companies already signed up to deploy Vera chips in their respective clouds: Alibaba, ByteDance, Meta, Oracle, CoreWeave, Lambda, Nebius, and NScale.
But the plot gets more interesting when Nvidia talks about the custom core. Previously, Nvidia’s Grace CPUs used off-the-shelf Arm cores, like Arm’s Neoverse V2 (GB200/300), Neoverse V3 (AGX Thor), or Cortex X925 and A725 cores (GB10). With Vera, Nvidia moves to designing its own ARMv9.2-compatible core called Olympus. The source also raises a useful skepticism: it is “up for debate” how custom the chip really is, because Nvidia says Olympus builds heavily on existing Arm IP. Looking at the architecture block diagram, the source observes Olympus has a 10-wide decoder and dispatch, eight integer ALUs, six vector/FP pipelines (SVE 128), four load units, and two store units. It looks, in many ways, like a heavily modified Cortex X925.
So where does the differentiation show up? Nvidia claims an entirely custom neural branch predictor that can explore two branches simultaneously. Branch prediction is performance-critical because CPUs try to guess which code path comes next and execute accordingly. By predicting two paths per cycle, Nvidia says Vera decreases misprediction likelihood. The source ties that directly to workloads like branch-heavy Python scripts generated on the fly by AI code assistants. Nvidia also claims core-level changes to reduce pipeline and execution bottlenecks that otherwise create pipeline stalls, aiming to maximize instructions per clock (IPC). Examples called out in the source include memory renaming to speed up store-to-load dependency chains, using it as a lever for pointer-heavy software, graph traversal, runtime frameworks, and complex object-oriented workloads. Nvidia also implements a value prediction scheme that “identifies stable dependency chains and predicts future values before they are produced,” allowing speculative execution while correctness is verified later. The company claims this helps repetitive software patterns and sequential data processing.
On the backend, Olympus includes eight ALUs for simple integer work and two complex ALUs that handle multiplication and division, CRC, and shift intensive workloads, plus four dedicated branch units to support control flow decisions. For floating point, Olympus features six 128-bit Arm scalable vector engines (SVE2) extensions with FP8 support, along with a pair of vector units dedicated to speeding up crypto operations. Memory hierarchy details are equally specific: 64 KB L1 instruction cache, 96 KB L1 data cache, 2MB L2 per core, and 164 MB system level cache (SLC) sharded across the chip.
One more “only-in-Arm-datacenters” twist: SMT. The source says Olympus supports simultaneous multithreading, but not in the conventional x86 SMT sense. Nvidia markets it as spatial multithreading, and the source clarifies that it is not simultaneous multi-threading where two threads share resources of one core. Instead, Olympus’ SMT implementation is described as “SMT-like,” with Nvidia using a different mechanism. That matters strategically because it hints how Nvidia plans to extract more utilization from the CPU in agent-host workloads, where different execution streams can create idle cycles.
The executive takeaway is straightforward: Vera is architected around removing the CPU bottlenecks that keep agent execution from stalling, while also coordinating GPU-heavy Vera Rubin systems. If the market believes the premise that agent hosting is a CPU job, Nvidia’s push does not just compete on performance. It changes purchasing behavior. Buyers could end up choosing Vera as a control-plane and execution-plane foundation, then layering GPUs as needed, rather than treating GPUs as the only entry point. For Intel, AMD, and every hyperscaler CFO negotiating capex, that is the real stakes.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

Samsung puts silicon carbon batteries into the new Fold, joining China’s adoption wave
See why silicon carbon batteries are spreading fast in phones, and what it signals for smartphone battery life bets.

NYPL saw teen programming attendance jump 27% since 2023 by designing curiosity-driven “third spaces”
Loneliness and screen fatigue are pushing young adults toward libraries, chess clubs, and book bars that make learning the social glue.

Stop settling for phone photos: third-party apps and editing can rival cameras
Smartphone sensors are getting better, but you can unlock near-camera results with extra tools and time.

