viable/strict/1791282642: [FSDP] Reduce TestParityWithDDP GPU delays 250ms -> 50ms (#199769)
- PyTorch: 1964 events in the last 90 days
- PyTorch: 1938th Release in the last 90 days
- Previous: earlier the same day · trunk/381cc3e476bad355609f87c1634e843f8d335df0: [audio hash update] update the pinned audio hash (#199869)
What happened
TestParityWithDDP's delay tests ( test_delayed_optim_step , test_delayed_reduce_scatter , test_mixture_of_experts_with_delay_before_free ) inject a 250ms torch.cuda._sleep per iteration so the GPU lags the CPU and missing FSDP stream syncs show up as parity mismatches. About 410s of the class's ~917s per CI config is these sleeps. TestParityWithDDPCUDA is the largest class in test_fsdp_core (~424 aggregate CI minute…
Summary assembled by rule from the sources below