Projects/H200 MoE Grouped GEMM profiling and tuning
GPU Performance · MoE · SM90

H200 MoE Grouped GEMM profiling and tuning

The main lesson was methodological: define the real workload before searching for a faster kernel.

4–6×Workload overcount corrected
+5.3%Target shape
+1.84%Full matrix aggregate
≈66.5%Pure GEMM MFU

Workload definition

Expanded rows are not compact rows

Mixing T × top-k expanded rows with DeepEP compact rows overstated work by roughly 4–6×. The corrected compact-row count changed both the benchmark distribution and the optimization target.

Replay

Production shapes, local control

Steady-state forward ranges were collected, filtered, and replayed locally. This preserved representative shapes while making kernel configuration, warmup, and repeated A/B measurement controllable.

Tuning

Configuration-level search

Candidate SM90 configurations varied tile shape, cluster shape, swizzle, and expert distribution. A selected target shape improved by about 5.3%, while the full shape/distribution matrix improved by about 1.84% in aggregate.

Claim boundary

Microkernel is not end-to-end

Pure GEMM MFU and local latency do not equal end-to-end MoE or model throughput. Dispatch, communication, imbalance, launch overhead, and synchronization remain outside that metric.