trunk/0993a6c66933eecdcfddbdc79a80f702067bc367: [Pipeline] Defer FSDP gradient reduction waits (#196725)
- PyTorch: 1673 events in the last 90 days
- PyTorch: 1654th Release in the last 90 days
- Previous: earlier the same day · trunk/769c18b73de9dd328fd434728589fffdf47846e0: [rpc] Fix dangling this in RRef::handleError static map (#192960)
What happened
Split pipeline FSDP gradient reduction into explicit start and wait actions. REDUCE_GRAD starts asynchronous FSDP gradient finalization and stores its public GradientReductionHandle . WAIT_REDUCE_GRAD waits for that handle and completes gradient scaling. Every reduction has exactly one matching wait, and each pipeline rank may have at most one pending reduction. By default, the wait remains adjacent to its reduction…
Summary assembled by rule from the sources below