← 返回事件
持续讨论AI发版

trunk/0993a6c66933eecdcfddbdc79a80f702067bc367: [Pipeline] Defer FSDP gradient reduction waits (#196725)

图:PyTorch Releases

发生了什么

Split pipeline FSDP gradient reduction into explicit start and wait actions. REDUCE_GRAD starts asynchronous FSDP gradient finalization and stores its public GradientReductionHandle . WAIT_REDUCE_GRAD waits for that handle and completes gradient scaling. Every reduction has exactly one matching wait, and each pipeline rank may have at most one pending reduction. By default, the wait remains adjacent to its reduction…

摘要按规则整理自下方来源原文

为什么在扩散

来源