trunk/1b51ed6169475ed863f4ee43b7eaa52ecda0f862: Register fused AdamW Meta kernel for tensor learning rates (#200362)
- PyTorch: 2113 events in the last 90 days
- PyTorch: 2084th Release in the last 90 days
- Previous: earlier the same day · ciflow/inductor/198050
What happened
Linked issue or supporting maintainer @sanketpurandare @tianyu-l @weifengpy Summary (human written only) To enable optimizer CUDA graph capture, we need to change learning rate to be a tensor. However, with fsdp2, the optimizer states are DTensor, which cause DTensor sharding propagation materializing the GPU tensors. The root cause is that there is no meta registration for optimizer with learning rate being a tenso…
Summary assembled by rule from the sources below