rl / Engineering

R3: MoE training–rollout route consistency

Connect vLLM route capture to Megatron replay, with development validation at 128 GPUs.

Problem & contribution

Natural routing differed between vLLM rollout and Megatron training. Route data, response masks and probability-drift metrics needed a shared contract.

Implemented route capture, transport and replay, with response-mask alignment and mismatch / fτ² / KL monitoring. Extended the rollout system to larger development models and GPU counts.

Validation

Condition / metricResult
8×H200 · Qwen3-30B-A3B · 20-step route mismatch17–19% → 0
Probability-drift metrics in that control experimentfτ² ↓36–145× · KL ↓4–7×
Separate large-scale development validation128 GPUs · 300B-class · 200-step rollout

Connected RL training and rollout through a tested data contract.

Source: the author’s September 2026 resume and engineering summary; results apply to the stated test conditions.

The 20-step smaller-model comparison and the 128-GPU, 300B-class 200-step rollout are separate validations. Route consistency alone does not establish improved reward, convergence or model quality.