Technical

Comparing Nvidia Data Center GPUs

Measured compute, memory bandwidth, and training throughput on four Nvidia GPUs, and what the differences mean in practice.

We benchmarked four Nvidia data center GPUs to see how they differ in compute, memory bandwidth, and training speed.

GPU Architecture Memory
A100 Ampere 40 GB HBM2
H100 Hopper 80 GB HBM3
H200 Hopper 141 GB HBM3e
RTX PRO 6000 Blackwell 96 GB GDDR7

How we measured

We ran one PyTorch script (PyTorch 2.8, CUDA 12.8) on a single GPU of each type, three times, and report the mean.

  • Compute: a 16384 x 16384 matrix multiply in TF32, FP16, BF16, and FP8. FP4 was measured separately on the RTX PRO 6000 with torchao and torch.compile.
  • Memory bandwidth: a 4 GiB device-to-device copy.
  • Training throughput: a GPT-style transformer with about 0.94 billion parameters, trained in bfloat16 and sized to fit the 40 GB A100.

These are achieved single-GPU numbers, not vendor peaks, and they do not include multi-GPU scaling.

Results

Switch between metrics and precisions to compare the four GPUs:

A100
267
H100
7102.7× A100
H200
6962.6× A100
RTX PRO 6000
3861.4× A100
Matrix multiply (BF16) in TFLOPS. Single GPU, mean of three runs.
GPU TF32 FP16 BF16 FP8 FP4 Bandwidth (GB/s) Training (tokens/s)
A100 124 263 267 n/a n/a 1,384 30,191
H100 355 690 710 1,339 n/a 3,042 73,899
H200 350 668 696 1,353 n/a 4,300 77,163
RTX PRO 6000 201 304 386 697 749 1,460 41,370

TF32 through FP8 are matrix multiply throughput in TFLOPS. FP4 uses the torchao NVFP4 path with torch.compile, so it is not directly comparable to the other columns.

Key findings

  • The H100 and H200 match on compute. They share the same compute architecture, and their BF16 results agree within run-to-run variation: 710 ± 14 TFLOPS on the H100 and 696 ± 5 on the H200. At BF16 both are about 2.6 to 2.7 times the A100 and 1.8 times the RTX PRO 6000.
  • FP8 roughly doubles matrix multiply throughput. It ran 1.8 to 1.9 times faster than BF16 on the H100, H200, and RTX PRO 6000. The A100 does not support FP8.
  • The H200 leads on memory bandwidth. At about 4,300 GB/s it is roughly 40 percent above the H100 and about three times the A100 and RTX PRO 6000.
  • Training tracks bandwidth, not just compute. The H200 trains this model about 4 percent faster than the H100, because the step is partly memory bound. Both Hopper GPUs are roughly 2.5 times faster than the A100, and the RTX PRO 6000 is about 37 percent faster than the A100.
  • FP4 is still early. Only the RTX PRO 6000 supports it. It reached about 749 TFLOPS, just above that GPU’s FP8, while a plain FP4 matrix multiply was several times slower, which points to software that is still maturing.
  • Memory capacity often decides whether a model fits at all: 40 GB on the A100, 80 GB on the H100, 141 GB on the H200, and 96 GB on the RTX PRO 6000.

Which GPU to pick

Pick a workload and set how much memory it needs:

What are you running?

A10040 GB · Ampere30 of 40 GB used
NVLink
H10080 GB · Hopper30 of 80 GB used
FP8NVLink
Best fitH200141 GB · Hopper30 of 141 GB used
FP8NVLink
RTX PRO 600096 GB · Blackwell30 of 96 GB used
FP8FP4RT CoresPCIe
Largest models, long context, memory-bound inferenceBest fit: H200. Most memory (141 GB) and highest bandwidth.

Reference

This post is adapted from a technical blog post I wrote with Yasin Mazloumi for our computing handbook. The original post includes the full benchmark script and error bars for every measurement.

All posts