Deterministic MoE router GEMM
Three backends for small-M decode with CUDA Graph preflight and bitwise checks.
Problem & contribution
Small-M decode wasted GEMM work while kernel selection and CUDA Graph execution had to satisfy determinism constraints.
Integrated DeepGEMM, Triton Full-K and persistent backends into vLLM. Implemented dispatch, Graph preflight, cache/fallback and workspace management, then validated kernel and model scenarios.
Validation
| Condition / metric | Result |
|---|---|
| Median latency in the measured 48-layer router workload | −73.39% |
| Bitwise-matching kernel cases | 135 / 135 |
| Bitwise-matching model scenarios | 20 / 20 |
| Full-request performance with prefix-cache hits | TP1 +3.81% · TP2 +4.36% |
Delivered the integration from kernel dispatch to inference runtime.
Source: the author’s September 2026 resume and engineering summary; results apply to the stated test conditions.
The 48 layers describe this measurement workload, not the released IQuest-Q1 architecture. The router-only reduction is not an end-to-end gain; request-level results apply to the tested prefix-cache-hit configurations.