inference / Engineering

Deterministic MoE router GEMM

Three backends for small-M decode with CUDA Graph preflight and bitwise checks.

Problem & contribution

Small-M decode wasted GEMM work while kernel selection and CUDA Graph execution had to satisfy determinism constraints.

Integrated DeepGEMM, Triton Full-K and persistent backends into vLLM. Implemented dispatch, Graph preflight, cache/fallback and workspace management, then validated kernel and model scenarios.

Validation

Condition / metricResult
Median latency in the measured 48-layer router workload−73.39%
Bitwise-matching kernel cases135 / 135
Bitwise-matching model scenarios20 / 20
Full-request performance with prefix-cache hitsTP1 +3.81% · TP2 +4.36%

Delivered the integration from kernel dispatch to inference runtime.

Source: the author’s September 2026 resume and engineering summary; results apply to the stated test conditions.

The 48 layers describe this measurement workload, not the released IQuest-Q1 architecture. The router-only reduction is not an end-to-end gain; request-level results apply to the tested prefix-cache-hit configurations.