Résumé

Haoran Feng

LLM systems & GPU performance engineering

Education

The Chinese University of Hong Kong, ShenzhenM.Sc. in Integrated Circuits and Systems · expected June 2027
Shandong UniversityB.Eng. in Electronic Science and Technology · Minor in Computer Science and Technology

Shandong University Academic Scholarship (top 20%) and Competition & Innovation Scholarship.

Experience

Ubiquant · IQuestLab

LLM & High-Performance Computing Intern

IQuest-Q1 · Contributed to training, inference and RL rollout infrastructure for the released IQuest-Q1 (320B-A15B) MoE model.

  • FA3 deterministic SWA backward · Isolated dQ semaphore dependencies and aligned contributor-relative tickets with the reverse scheduler. Full backward on a representative packed shape: 6.755 → 1.699 ms (3.98×); 3,601 regressions and 1,000 bitwise checks; wheel delivered.
  • FC1 compute–communication · Four-H200 dual-stream proxy with NCCL CTA / DeepGEMM SM budgeting; 121 coarse and 89 fine configurations plus rechecks. AllGatherV proxy joint latency −12.80%; isolated compute-budget contribution 7.58%.
  • Deterministic router GEMM · Three backends, Graph preflight, cache/fallback and workspace management. Median latency −73.39% in a 48-layer test workload; 135/135 kernel and 20/20 model cases bitwise-matching; full-request performance +3.81% / +4.36% at TP1 / TP2 with prefix-cache hits.
  • OE asynchronous state · GPU token history and Triton fused hashing. TP1: 28/28 cases, throughput +4.99%. TP2 / TP4 safe paths: 84/84 each, zero mismatch, decode throughput +4.6% / +3.0%.
  • R3 Router Replay · Connected vLLM / Megatron route capture, transport, replay and monitoring. Eight-H200, Qwen3-30B-A3B, 20-step control: mismatch 17–19% → 0, fτ² / KL lower by 36–145× / 4–7×. Separately completed a 200-step development rollout for a 300B-class model on 128 GPUs.
  • CUDA Graph stability · A two-GPU reproducer isolated FTZ-sensitive Lamport sentinel handling. Backported the upstream bitwise fix; repeated model-level Graph replay showed no further hang or timeout.

Moore Threads

Operator & Compiler Optimization Intern

TensorFlow MUSA Extension

  • Integrated muDNN and GELU fusion: 11/11 GELUs fused; operator latency −36.6% on workload-shaped tests. Logical_Or: 21.2 → 10.7 μs. Implemented or fixed Adam / AdaMax / GradientDescent and tests.
  • Reworked HostMemory and resource/session lifetimes. Stability improved from approximately 30% success over 500-round runs to 100% over 1,000-round runs, followed by 40k / 400k / 800k-round validation.

Open source

All PR links
ProjectContextMerged
MooncakeCommunity7
TensorFlow MUSADuring internship10
RayCommunity1
FlashInferCommunity1
MirageCommunity1

20 merged PRs across five AI infrastructure projects, including 10 from the Moore Threads internship. Verified 2026-09-29.

Projects & research

  • Adapted W4A16 / W8A8 / FP8 backends and integrated FlashAttention-2/4, GQA, softcap and SDPA. Used Compute Sanitizer to isolate gemv_kernel_g128 tail-group out-of-bounds accesses; added guards, packed-AWQ TP=2 column/row sharding and NCCL all-reduce.
  • Paired Qwen2.5-Coder-32B W4A16 deployment: checkpoint / peak memory −70.5% / −66.8%; output throughput +76.8%. Four sanitizer checks on representative shapes showed zero errors/hazards; numerical, cross-rank token and interactive acceptance passed.
AAAI 2027 · Submitted

From Structure to Preference: Token Weighting for Chain-of-Thought Distillation in Large Language Models

Shankui Han, Zhaoyu Li, Weiwen Yuan, Yuyuan Yang, Haoran Feng, Jinke Ren

Co-author. Responsible for model training, hyperparameter tuning and vLLM TP=4 evaluation, including LoRA / DPO training and controlled comparisons.

Contributed to a three-stage framework of structure recovery, key-token-weighted supervision and preference optimization. Token importance is estimated by the drop in teacher likelihood of the reference answer after perturbing each rationale token, then used in weighted SFT and an auxiliary loss.

Student / comparisonGSM8KSVAMP
Qwen2.5-7B-Instruct94.01%94.00%
Gain over strongest baseline+5.51 pp+10.10 pp

OpenReview ↗ · Results are reported in the submitted manuscript. Submission does not imply acceptance.

Skills

C/C++, Python, CUDA, Triton, PyTorch Extension; Nsys, NCU, Compute Sanitizer; attention, MoE, quantized GEMM, tensor parallelism and CUDA Graph.

IELTS 6.5 · CET-6 · HCIA-AI