Skip to main content
The NVIDIA Hopper architecture (compute capability 9.0) represents a major leap in GPU computing with Warpgroup Matrix Multiply-Accumulate (WGMMA), Tensor Memory Accelerator (TMA), and native FP8 support.

Supported GPUs

Hopper architecture-accelerated features require compiling with the sm_90a target (note the “a” suffix) to enable WGMMA and TMA instructions:

Key Features

1. Warpgroup Matrix Multiply-Accumulate (WGMMA)

WGMMA operates on 4 warps (128 threads) simultaneously, dramatically improving instruction throughput:
WGMMA Synchronization Primitives:
Usage Pattern:

2. Tensor Memory Accelerator (TMA)

TMA provides hardware-accelerated bulk data transfer between global and shared memory:
TMA Benefits:
  • Hardware-managed address generation
  • Automatic boundary checking
  • Integrated with async barriers
  • Support for up to 5D tensors
  • Prefetch capabilities

3. Enhanced FP64 Tensor Cores

Hopper provides improved FP64 tensor core shapes:
Supported FP64 Shapes:
  • 16x8x4 (new in Hopper)
  • 16x8x8 (new in Hopper)
  • 16x8x16 (new in Hopper)

4. FP8 Tensor Core Support

Hopper introduces native FP8 support with two formats:
  • E4M3: 4 exponent bits, 3 mantissa bits (better for forward pass)
  • E5M2: 5 exponent bits, 2 mantissa bits (better for gradients)

Instruction Shapes

Notice the significantly larger M and K dimensions compared to Ampere’s 16x8xK shapes. This enables more efficient computation per instruction.

Thread Block Clusters

Hopper introduces thread block clusters for multi-CTA cooperation:
Cluster Configuration:

Asynchronous Barriers

Hopper enhances barrier functionality for TMA integration:

Complete WGMMA+TMA Example

Performance Optimization

Optimal Tile Sizes for WGMMA

Cluster Configuration

  • Small problems: 1x1x1 cluster (single CTA)
  • Medium problems: 2x1x1 or 1x2x1 cluster
  • Large problems: 2x2x1 cluster (maximum practical)

Pipeline Stages

  • WGMMA kernels: 3-4 stages typical
  • Balance: TMA latency vs shared memory capacity
  • Hopper H100: Up to 227 KB shared memory per SM

Compilation

Using sm_90 (without “a”) will not enable WGMMA or TMA. Always use sm_90a for Hopper architecture-accelerated features.

Examples

Hopper-specific examples in CUTLASS:
  • examples/48_hopper_warp_specialized_gemm/ - WGMMA-based GEMM
  • examples/57_hopper_warp_specialized_gemm_with_epilogue_visitor/ - Epilogue fusion
  • examples/58_hopper_persistent_gemm/ - Persistent kernel design
  • examples/111_hopper_ssd/ - State Space Decomposition

See Also