Server chips have traditionally lagged client parts in single-thread performance, but NVIDIA's Olympus core aims to close that gap. Instead of chasing high clock speeds, Olympus runs at a modest 3.3 GHz and focuses on performance per clock, using a 10-wide out-of-order design with unusually large structures and simultaneous multithreading.
According to a detailed analysis by Chips and Cheese, Olympus shares a high-level layout with Arm's Cortex X925 but pushes out-of-order structure sizes further. The branch predictor is a key strength: in SPEC CPU2026 integer tests it lands slightly behind AMD's Zen 5 and slightly ahead of Intel's Lion Cove, while in floating-point it comes out ahead, helped by a large win on 731.astcenc.
The core can handle two taken branches per cycle, similar to Zen 5, though its capability is tied to instruction footprint rather than branch count. NVIDIA also uses ahead indexing to improve predictor throughput. The result, the article says, is a "monster core" that gets very close to desktop performance in the server scene.