FC1: compute–communication contention
Jointly tune NCCL CTA and DeepGEMM SM budgets in a dual-stream, four-H200 proxy.
Problem & contribution
Communication on a separate stream competed with FC1 for GPU resources, degrading concurrent compute performance.
Built a dual-stream proxy on four H200 GPUs and co-tuned NCCL CTA and DeepGEMM SM budgets. Ran 121 coarse configurations, 89 fine configurations and rechecks; held communication settings fixed to isolate the compute-budget contribution.
Validation
| Condition / metric | Result |
|---|---|
| AllGatherV proxy joint latency | −12.80% |
| Compute-budget contribution at fixed communication settings | 7.58% |
| Configuration sweeps | 121 + 89 |
Produced a joint-budget selection and revalidation procedure.
Source: the author’s September 2026 resume and engineering summary; results apply to the stated test conditions.
These are dual-stream proxy results, not full-model training throughput. Joint gains and isolated compute-budget gains use different controls and must not be added.