trunk/68d62895fad677a8497f83eef2e0e348d7c7f1ab: [cuda] Compact ragged partials for foreach norm/max reductions (#190549)
- PyTorch: 1835 events in the last 90 days
- PyTorch: 1811th Release in the last 90 days
- Previous: earlier the same day · trunk/8494d3b14dbe5a738fe12ecf5cf0163fc4a2b99f: [docs] Fix config option in PGO docs (#191027)
What happened
Summary The CUDA _foreach_norm and _foreach_max kernels reduce each tensor in two steps: a first kernel writes one partial result per fixed-size chunk into a scratch buffer, and a cleanup kernel reduces each tensor's partials into its final value. _foreach_norm is what torch.nn.utils.clip_grad_norm_ and the fused optimizers call every step, so this path is hot during training. The scratch buffer was allocated as a r…
Summary assembled by rule from the sources below