The NVIDIA H200 is a Hopper-architecture data center GPU and the first NVIDIA GPU to feature HBM3e memory. It provides 141 GB of HBM3e with substantially higher memory bandwidth than the H100 (approximately 4.8 TB/s), making it especially well suited to memory-bound workloads.
- 141 GB HBM3e — nearly double the H100 80GB, allowing larger models, longer context, and bigger batch sizes in a single GPU.
- ~4.8 TB/s memory bandwidth to keep Tensor Cores fed on bandwidth-limited operations.
- Same proven Hopper compute platform as the H100 (FP8 Transformer Engine), so CUDA, cuDNN, and TensorRT-LLM stacks carry over.
- Available in SXM and other data center form factors with NVLink multi-GPU scaling.
Typical use cases: memory-bound LLM inference, long-context serving, high-batch inference, retrieval-augmented generation, and memory-intensive HPC and training.
A natural upgrade path for many H100 deployments, with the biggest gains on memory-bound inference. Confirm system/baseboard compatibility and power/cooling before ordering. Contact us for availability.




Reviews
There are no reviews yet.