← 返回事件
持续讨论AI发版

trunk/68d62895fad677a8497f83eef2e0e348d7c7f1ab: [cuda] Compact ragged partials for foreach norm/max reductions (#190549)

图:PyTorch Releases

发生了什么

Summary The CUDA _foreach_norm and _foreach_max kernels reduce each tensor in two steps: a first kernel writes one partial result per fixed-size chunk into a scratch buffer, and a cleanup kernel reduces each tensor's partials into its final value. _foreach_norm is what torch.nn.utils.clip_grad_norm_ and the fused optimizers call every step, so this path is hot during training. The scratch buffer was allocated as a r…

摘要按规则整理自下方来源原文

为什么在扩散

来源