- Nvidia’s Olympus core architecture prioritizes single-threaded IPC over frequency using a 10-wide decoding frontend, deep out-of-order midcore, and a graph prefetcher tuned for agent AI’s branch-heavy, pointer-heavy workloads
- Vera trades chiplet-style core density for a monolithic 88-core die and single-NUMA-per-socket design, a deliberate bet on agentic AI workloads that Nvidia’s own engineers admit comes at the expense of legacy workload performance
- Self-reported SPEC CPU 2026 results by Nvidia draw a per-core advantage figure of somewhere between 70% and 80% over AMD’s EPYC 9755 server CPU
Nvidia’s Vera CPU has a lot to prove, representing the company’s first real foray into the AI server CPU market, where traditional vendors Intel and AMD, along with third-party Arm-based providers, are all seeking a piece of an increasingly lucrative data center pie.
To that end, Nvidia has published its most detailed technical report yet on the chip, promising significant performance gains over the competition. Vera is the company’s first server processor built around a fully in-house core design, a departure from the regular Arm cores that powered its Grace predecessor.
While Vera is headed for general availability in the second half of 2026, Nvidia is making its case on a specific front: not raw core count, but sustained per-core performance under load, the metric it claims matters most for agent AI.
Latest videos fromTechRadar
What’s really under the hood of Nvidia’s Vera CPU?
At the heart of each Vera CPU is Olympus, Nvidia’s first custom server core, built for the Armv9.2 instruction set, but designed in-house instead of being derived from Arm’s standard Neoverse design, as its predecessor, Grace, was.
Nvidia’s second-generation data center CPU is a completely redesigned design built around a single goal: to lead in agent workload performance.
Instead of following Intel and AMD down the chiplet route, Nvidia packs all 88 cores (176 threads) onto a single monolithic die, with a dual-socket configuration that delivers 176 cores and 352 threads in one system.
Nvidia has detailed the core design extensively, and several publications have since dug into the microarchitecture. The front end runs a 10-wide decoding engine paired with a neural branch predictor that can resolve up to two branches per cycle, designed to handle the large instruction footprints and irregular control flow of interpreters, compilers, and agent runtimes.
The mid-core combines a broad renaming and allocation engine with a large reorder buffer and dependency-breaking techniques, including memory renaming and value prediction.
The execution engine dynamically schedules across integer, vector, floating-point, cryptographic, load, and storage resources, while the cache subsystem adds a graph prefetcher that targets the pointer-chasing access patterns that defeat conventional streaming prefetchers.
That design has not gone unanswered. AMD has countered with estimated numbers for its 256-core Zen 6 “Venice” part, claiming a 3.3x rack-level advantage over Vera, though those numbers are extrapolated rather than measured.
AMD reacting first is not surprising. In x86, it continues to take ground from Intel, reaching 33.2% of x86 server CPU shipments in Q1 2026 per Mercury Research, up from 27.2% the year before.
With Arm reporting that its architecture now accounts for about 50% of CPU compute among the top hyperscalers, the competition is building from several directions at once. Vera may well set a new bar for agent AI.
Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews and opinions in your feeds.



