training / Engineering

FC1: compute–communication contention

Jointly tune NCCL CTA and DeepGEMM SM budgets in a dual-stream, four-H200 proxy.

Problem & contribution

Communication on a separate stream competed with FC1 for GPU resources, degrading concurrent compute performance.

Built a dual-stream proxy on four H200 GPUs and co-tuned NCCL CTA and DeepGEMM SM budgets. Ran 121 coarse configurations, 89 fine configurations and rechecks; held communication settings fixed to isolate the compute-budget contribution.

Validation

Condition / metricResult
AllGatherV proxy joint latency−12.80%
Compute-budget contribution at fixed communication settings7.58%
Configuration sweeps121 + 89

Produced a joint-budget selection and revalidation procedure.

Source: the author’s September 2026 resume and engineering summary; results apply to the stated test conditions.

These are dual-stream proxy results, not full-model training throughput. Joint gains and isolated compute-budget gains use different controls and must not be added.