R3: MoE training–rollout route consistency
Connect vLLM route capture to Megatron replay, with development validation at 128 GPUs.
Problem & contribution
Natural routing differed between vLLM rollout and Megatron training. Route data, response masks and probability-drift metrics needed a shared contract.
Implemented route capture, transport and replay, with response-mask alignment and mismatch / fτ² / KL monitoring. Extended the rollout system to larger development models and GPU counts.
Validation
| Condition / metric | Result |
|---|---|
| 8×H200 · Qwen3-30B-A3B · 20-step route mismatch | 17–19% → 0 |
| Probability-drift metrics in that control experiment | fτ² ↓36–145× · KL ↓4–7× |
| Separate large-scale development validation | 128 GPUs · 300B-class · 200-step rollout |
Connected RL training and rollout through a tested data contract.
Source: the author’s September 2026 resume and engineering summary; results apply to the stated test conditions.
The 20-step smaller-model comparison and the 128-GPU, 300B-class 200-step rollout are separate validations. Route consistency alone does not establish improved reward, convergence or model quality.