Comparing Nvidia Data Center GPUs
Measured compute, memory bandwidth, and training throughput on four Nvidia GPUs, and what the differences mean in practice.
We benchmarked four Nvidia data center GPUs to see how they differ in compute, memory bandwidth, and training speed.
| GPU | Architecture | Memory |
|---|---|---|
| A100 | Ampere | 40 GB HBM2 |
| H100 | Hopper | 80 GB HBM3 |
| H200 | Hopper | 141 GB HBM3e |
| RTX PRO 6000 | Blackwell | 96 GB GDDR7 |
How we measured
We ran one PyTorch script (PyTorch 2.8, CUDA 12.8) on a single GPU of each type, three times, and report the mean.
- Compute: a 16384 x 16384 matrix multiply in TF32, FP16, BF16, and FP8. FP4 was measured separately on the RTX PRO 6000 with torchao and torch.compile.
- Memory bandwidth: a 4 GiB device-to-device copy.
- Training throughput: a GPT-style transformer with about 0.94 billion parameters, trained in bfloat16 and sized to fit the 40 GB A100.
These are achieved single-GPU numbers, not vendor peaks, and they do not include multi-GPU scaling.
Results
Switch between metrics and precisions to compare the four GPUs:
| GPU | TF32 | FP16 | BF16 | FP8 | FP4 | Bandwidth (GB/s) | Training (tokens/s) |
|---|---|---|---|---|---|---|---|
| A100 | 124 | 263 | 267 | n/a | n/a | 1,384 | 30,191 |
| H100 | 355 | 690 | 710 | 1,339 | n/a | 3,042 | 73,899 |
| H200 | 350 | 668 | 696 | 1,353 | n/a | 4,300 | 77,163 |
| RTX PRO 6000 | 201 | 304 | 386 | 697 | 749 | 1,460 | 41,370 |
TF32 through FP8 are matrix multiply throughput in TFLOPS. FP4 uses the torchao NVFP4 path with torch.compile, so it is not directly comparable to the other columns.
Key findings
- The H100 and H200 match on compute. They share the same compute architecture, and their BF16 results agree within run-to-run variation: 710 ± 14 TFLOPS on the H100 and 696 ± 5 on the H200. At BF16 both are about 2.6 to 2.7 times the A100 and 1.8 times the RTX PRO 6000.
- FP8 roughly doubles matrix multiply throughput. It ran 1.8 to 1.9 times faster than BF16 on the H100, H200, and RTX PRO 6000. The A100 does not support FP8.
- The H200 leads on memory bandwidth. At about 4,300 GB/s it is roughly 40 percent above the H100 and about three times the A100 and RTX PRO 6000.
- Training tracks bandwidth, not just compute. The H200 trains this model about 4 percent faster than the H100, because the step is partly memory bound. Both Hopper GPUs are roughly 2.5 times faster than the A100, and the RTX PRO 6000 is about 37 percent faster than the A100.
- FP4 is still early. Only the RTX PRO 6000 supports it. It reached about 749 TFLOPS, just above that GPU’s FP8, while a plain FP4 matrix multiply was several times slower, which points to software that is still maturing.
- Memory capacity often decides whether a model fits at all: 40 GB on the A100, 80 GB on the H100, 141 GB on the H200, and 96 GB on the RTX PRO 6000.
Which GPU to pick
Pick a workload and set how much memory it needs:
What are you running?
Reference
This post is adapted from a technical blog post I wrote with Yasin Mazloumi for our computing handbook. The original post includes the full benchmark script and error bars for every measurement.