trunk/68d62895fad677a8497f83eef2e0e348d7c7f1ab: [cuda] Compact ragged partials for foreach norm/max reductions (#190549)
- PyTorch 近 90 天出现 1839 次
- PyTorch 近 90 天第 1815 次发版
- 上一次:同一天稍早 · trunk/8494d3b14dbe5a738fe12ecf5cf0163fc4a2b99f: [docs] Fix config option in PGO docs (#191027)
发生了什么
Summary The CUDA _foreach_norm and _foreach_max kernels reduce each tensor in two steps: a first kernel writes one partial result per fixed-size chunk into a scratch buffer, and a cleanup kernel reduces each tensor's partials into its final value. _foreach_norm is what torch.nn.utils.clip_grad_norm_ and the fused optimizers call every step, so this path is hot during training. The scratch buffer was allocated as a r…
摘要按规则整理自下方来源原文