H200 MoE Grouped GEMM profiling and tuning
The main lesson was methodological: define the real workload before searching for a faster kernel.
Workload definition
Expanded rows are not compact rowsMixing T × top-k expanded rows with DeepEP compact rows overstated work by roughly 4–6×. The corrected compact-row count changed both the benchmark distribution and the optimization target.
Replay
Production shapes, local controlSteady-state forward ranges were collected, filtered, and replayed locally. This preserved representative shapes while making kernel configuration, warmup, and repeated A/B measurement controllable.
Tuning
Configuration-level searchCandidate SM90 configurations varied tile shape, cluster shape, swizzle, and expert distribution. A selected target shape improved by about 5.3%, while the full shape/distribution matrix improved by about 1.84% in aggregate.
Claim boundary
Microkernel is not end-to-endPure GEMM MFU and local latency do not equal end-to-end MoE or model throughput. Dispatch, communication, imbalance, launch overhead, and synchronization remain outside that metric.