Ubiquant · IQuestLab · April 2026–present
IQuest-Q1: model infrastructure
Contributed to IQuest-Q1 training, inference and RL rollout infrastructure through kernel optimization, training–rollout consistency, stability debugging and multi-GPU validation.
Team release
320B / 15BTotal / active parameters per token
88 layersReleased model depth
512KReleased context length
The official release describes IQuest-Q1 as a MoE model for agentic coding, reasoning and multi-step tool use, with a 3 SWA + 1 FA attention pattern. These specifications describe the team’s released model.
My contribution
01
Training kernels
Deterministic SWA backward; FC1 compute–communication contention.
02
Inference runtime
Deterministic router GEMM; asynchronous GPU token state.
03
RL infrastructure
Router Replay; multi-GPU rollout and CUDA Graph stability.
01FA3: deterministic SWA backward schedulingShorten the dQ semaphore dependency chain inside the fused kernel.Full backward · 3.98×02FC1: compute–communication contentionJointly tune NCCL CTA and DeepGEMM SM budgets in a dual-stream, four-H200 proxy.Proxy joint latency −12.80%03Deterministic MoE router GEMMThree backends for small-M decode with CUDA Graph preflight and bitwise checks.Router latency −73.39%04OE: asynchronous GPU state operatorsKeep token history on the GPU and reduce synchronization with Triton fused hashing.TP1 throughput +4.99%05R3: MoE training–rollout route consistencyConnect vLLM route capture to Megatron replay, with development validation at 128 GPUs.Route mismatch → 006FlashInfer: diagnosing a CUDA Graph hangA two-GPU reproducer isolated FTZ-sensitive sentinel handling; an upstream fix was backported.Model-level Graph regression