Ubiquant · IQuestLab
LLM & High-Performance Computing Intern
IQuest-Q1 · Contributed to training, inference and RL rollout infrastructure for the released IQuest-Q1 (320B-A15B) MoE model.
- FA3 deterministic SWA backward · Isolated dQ semaphore dependencies and aligned contributor-relative tickets with the reverse scheduler. Full backward on a representative packed shape: 6.755 → 1.699 ms (3.98×); 3,601 regressions and 1,000 bitwise checks; wheel delivered.
- FC1 compute–communication · Four-H200 dual-stream proxy with NCCL CTA / DeepGEMM SM budgeting; 121 coarse and 89 fine configurations plus rechecks. AllGatherV proxy joint latency −12.80%; isolated compute-budget contribution 7.58%.
- Deterministic router GEMM · Three backends, Graph preflight, cache/fallback and workspace management. Median latency −73.39% in a 48-layer test workload; 135/135 kernel and 20/20 model cases bitwise-matching; full-request performance +3.81% / +4.36% at TP1 / TP2 with prefix-cache hits.
- OE asynchronous state · GPU token history and Triton fused hashing. TP1: 28/28 cases, throughput +4.99%. TP2 / TP4 safe paths: 84/84 each, zero mismatch, decode throughput +4.6% / +3.0%.
- R3 Router Replay · Connected vLLM / Megatron route capture, transport, replay and monitoring. Eight-H200, Qwen3-30B-A3B, 20-step control: mismatch 17–19% → 0, fτ² / KL lower by 36–145× / 4–7×. Separately completed a 200-step development rollout for a 300B-class model on 128 GPUs.
- CUDA Graph stability · A two-GPU reproducer isolated FTZ-sensitive Lamport sentinel handling. Backported the upstream bitwise fix; repeated model-level Graph replay showed no further hang or timeout.