AI Infrastructure
Haoran Feng冯浩然
M.Sc. candidate at CUHK-Shenzhen, focused on LLM systems and GPU performance.
- Current
- IQuestLab · AI Infra Intern
IQuest-Q1 · Infrastructure contributions to a released model20 PRs · including internship workH200 / B200 · development access
Selected work
All work →01IQuest-Q1: training, inference and RL infrastructureContributed kernel optimization, training–rollout consistency and stability work to the 320B-A15B MoE model.Released model · internship02FA3: deterministic SWA backward schedulingShorten the dQ semaphore dependency chain inside the fused kernel.Full backward · 3.98×03Deterministic MoE router GEMMThree backends for small-M decode with CUDA Graph preflight and bitwise checks.Router latency −73.39%04LLMQRT: quantized inference runtimeA PyTorch Extension runtime for two-H200, 32B quantized inference and tensor parallelism.Output throughput +76.8%
Open source
PR record →| Project | Context | Merged |
|---|---|---|
| Mooncake | Community | 7 |
| TensorFlow MUSA | During internship | 10 |
| Ray | Community | 1 |
| FlashInfer | Community | 1 |
| Mirage | Community | 1 |
20 merged PRs across five AI infrastructure projects, including 10 from the Moore Threads internship. Verified 2026-09-29.
Education
The Chinese University of Hong Kong, ShenzhenM.Sc. in Integrated Circuits and Systems · expected June 2027
Shandong UniversityB.Eng. in Electronic Science and Technology · Minor in Computer Science and Technology
Research
AAAI 2027 · Submitted