Nvidia reveals deep dive into what makes Vera CPU tick — with the super-powered Olympus core promising huge performance increases and more

A different kind of AI data center CPU

by · TechRadar

News By Rahim Amir Published 24 July 2026

(Image credit: Nvidia)

Share this article 0 Join the conversation Follow us Add us as a preferred source on Google Newsletter Subscribe to our newsletter


  • Nvidia's Olympus core architecture prioritizes single-thread IPC over frequency, using a 10-wide decode front end, deep out-of-order mid-core, and a graph prefetcher tuned for agentic AI's branch-heavy, pointer-heavy workloads
  • Vera trades chiplet-style core density for a monolithic 88-core die and single-NUMA-per-socket design, a deliberate bet on agentic AI workloads that Nvidia's own engineers admit comes at the expense of legacy workload performance
  • Self-reported SPEC CPU 2026 results by Nvidia paint a per-core advantage figure of anywhere between 70% and 80% versus AMD's EPYC 9755 server CPU

Nvidia's Vera CPU has a lot to prove, representing the company's first real attempt at the AI server CPU market, where traditional vendors Intel and AMD, along with third-party Arm-based providers, are all gunning for a piece of an increasingly lucrative data center pie.

To that end, Nvidia has published its most detailed technical account yet of the chip, promising substantial performance gains over the competition. Vera is the company's first server processor built around a fully in-house core design, a departure from the stock Arm cores that powered its Grace predecessor.

As Vera heads toward general availability in the second half of 2026, Nvidia is making its case on a specific front: not raw core count, but sustained per-core performance under load, the metric it argues matters most for agentic AI.

Latest Videos FromTechRadarWatch full video here:

What's actually under the hood of Nvidia's Vera CPU?

At the core of every Vera CPU is Olympus, Nvidia's first custom server core, built to the Armv9.2 instruction set but designed in-house rather than derived from Arm's stock Neoverse designs, as its predecessor, Grace, was.

Nvidia's second-generation data center CPU is a completely reworked design built around a single goal: to lead in agentic workload performance.

Rather than following Intel and AMD down the chiplet route, Nvidia packs all 88 cores (176 threads) onto a single monolithic die, with a dual-socket configuration delivering 176 cores and 352 threads in one system.

Nvidia has detailed the core design extensively, and several publications have since dug into the microarchitecture. The front end runs a 10-wide decode engine paired with a neural branch predictor that can resolve up to two taken branches per cycle, designed to handle the large instruction footprints and irregular control flow of interpreters, compilers, and agent runtimes.

Are you a pro? Subscribe to our newsletter

Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!

Contact me with news and offers from other Future brandsReceive email from us on behalf of our trusted partners or sponsors