Projects
Engineering work
Model infrastructure, GPU kernels and inference runtime. Contributions first, experimental context next.
IQuest-Q1 · Internship engineering
Released model; view my contribution →
01FA3: deterministic SWA backward schedulingShorten the dQ semaphore dependency chain inside the fused kernel.Full backward · 3.98×02FC1: compute–communication contentionJointly tune NCCL CTA and DeepGEMM SM budgets in a dual-stream, four-H200 proxy.Proxy joint latency −12.80%03Deterministic MoE router GEMMThree backends for small-M decode with CUDA Graph preflight and bitwise checks.Router latency −73.39%04OE: asynchronous GPU state operatorsKeep token history on the GPU and reduce synchronization with Triton fused hashing.TP1 throughput +4.99%05R3: MoE training–rollout route consistencyConnect vLLM route capture to Megatron replay, with development validation at 128 GPUs.Route mismatch → 006FlashInfer: diagnosing a CUDA Graph hangA two-GPU reproducer isolated FTZ-sensitive sentinel handling; an upstream fix was backported.Model-level Graph regression
Runtime & frameworks
Other engineering records
| Project | Focus |
|---|---|
| H200 Grouped GEMM | Workload reconstruction, shape replay and tuning. |
| DeepGEMM offline delivery | Offline cubins, integrity checks and deployment. |
| Drone detection experiment pipeline | Data engineering, multi-GPU scheduling and training recovery. |