Supported GPUs
Key Features
1. Warpgroup Matrix Multiply-Accumulate (WGMMA)
WGMMA operates on 4 warps (128 threads) simultaneously, dramatically improving instruction throughput:2. Tensor Memory Accelerator (TMA)
TMA provides hardware-accelerated bulk data transfer between global and shared memory:- Hardware-managed address generation
- Automatic boundary checking
- Integrated with async barriers
- Support for up to 5D tensors
- Prefetch capabilities
3. Enhanced FP64 Tensor Cores
Hopper provides improved FP64 tensor core shapes:- 16x8x4 (new in Hopper)
- 16x8x8 (new in Hopper)
- 16x8x16 (new in Hopper)
4. FP8 Tensor Core Support
Hopper introduces native FP8 support with two formats:- E4M3: 4 exponent bits, 3 mantissa bits (better for forward pass)
- E5M2: 5 exponent bits, 2 mantissa bits (better for gradients)
Instruction Shapes
Notice the significantly larger M and K dimensions compared to Ampere’s 16x8xK shapes. This enables more efficient computation per instruction.
Thread Block Clusters
Hopper introduces thread block clusters for multi-CTA cooperation:Asynchronous Barriers
Hopper enhances barrier functionality for TMA integration:Complete WGMMA+TMA Example
Performance Optimization
Optimal Tile Sizes for WGMMA
Cluster Configuration
- Small problems: 1x1x1 cluster (single CTA)
- Medium problems: 2x1x1 or 1x2x1 cluster
- Large problems: 2x2x1 cluster (maximum practical)
Pipeline Stages
- WGMMA kernels: 3-4 stages typical
- Balance: TMA latency vs shared memory capacity
- Hopper H100: Up to 227 KB shared memory per SM
Compilation
Examples
Hopper-specific examples in CUTLASS:examples/48_hopper_warp_specialized_gemm/- WGMMA-based GEMMexamples/57_hopper_warp_specialized_gemm_with_epilogue_visitor/- Epilogue fusionexamples/58_hopper_persistent_gemm/- Persistent kernel designexamples/111_hopper_ssd/- State Space Decomposition
See Also
- Ampere Architecture - Previous generation
- Blackwell Architecture - Next generation
- CuTe Documentation
- NVIDIA Hopper Architecture Whitepaper