Education / Experience

Experience

From GPU backends to training and inference systems for foundation models.

Education

The Chinese University of Hong Kong, ShenzhenM.Sc. in Integrated Circuits and Systems · expected June 2027
Shandong UniversityB.Eng. in Electronic Science and Technology · Minor in Computer Science and Technology

Shandong University Academic Scholarship (top 20%) and Competition & Innovation Scholarship.

Experience

Ubiquant · IQuestLab

LLM & High-Performance Computing Intern

IQuest-Q1 · Contributed to training, inference and RL rollout infrastructure for the released IQuest-Q1 (320B-A15B) MoE model.

  • FA3 deterministic SWA backward · Isolated dQ semaphore dependencies and aligned contributor-relative tickets with the reverse scheduler. Full backward on a representative packed shape: 6.755 → 1.699 ms (3.98×); 3,601 regressions and 1,000 bitwise checks; wheel delivered.
  • FC1 compute–communication · Four-H200 dual-stream proxy with NCCL CTA / DeepGEMM SM budgeting; 121 coarse and 89 fine configurations plus rechecks. AllGatherV proxy joint latency −12.80%; isolated compute-budget contribution 7.58%.
  • Deterministic router GEMM · Three backends, Graph preflight, cache/fallback and workspace management. Median latency −73.39% in a 48-layer test workload; 135/135 kernel and 20/20 model cases bitwise-matching; full-request performance +3.81% / +4.36% at TP1 / TP2 with prefix-cache hits.
  • OE asynchronous state · GPU token history and Triton fused hashing. TP1: 28/28 cases, throughput +4.99%. TP2 / TP4 safe paths: 84/84 each, zero mismatch, decode throughput +4.6% / +3.0%.
  • R3 Router Replay · Connected vLLM / Megatron route capture, transport, replay and monitoring. Eight-H200, Qwen3-30B-A3B, 20-step control: mismatch 17–19% → 0, fτ² / KL lower by 36–145× / 4–7×. Separately completed a 200-step development rollout for a 300B-class model on 128 GPUs.
  • CUDA Graph stability · A two-GPU reproducer isolated FTZ-sensitive Lamport sentinel handling. Backported the upstream bitwise fix; repeated model-level Graph replay showed no further hang or timeout.

Moore Threads

Operator & Compiler Optimization Intern

TensorFlow MUSA Extension

  • Integrated muDNN and GELU fusion: 11/11 GELUs fused; operator latency −36.6% on workload-shaped tests. Logical_Or: 21.2 → 10.7 μs. Implemented or fixed Adam / AdaMax / GradientDescent and tests.
  • Reworked HostMemory and resource/session lifetimes. Stability improved from approximately 30% success over 500-round runs to 100% over 1,000-round runs, followed by 40k / 400k / 800k-round validation.

Research

AAAI 2027 · Submitted

Co-author; model training, hyperparameter tuning and vLLM TP=4 evaluation.

Public engineering record

View PRs →
ProjectContextMerged
MooncakeCommunity7
TensorFlow MUSADuring internship10
RayCommunity1
FlashInferCommunity1
MirageCommunity1

20 merged PRs across five AI infrastructure projects, including 10 from the Moore Threads internship. Verified 2026-09-29.