trunk/1b51ed6169475ed863f4ee43b7eaa52ecda0f862: Register fused AdamW Meta kernel for tensor learning rates (#200362)
- PyTorch 近 90 天出现 2115 次
- PyTorch 近 90 天第 2086 次发版
- 上一次:同一天稍早 · ciflow/inductor/198050
发生了什么
Linked issue or supporting maintainer @sanketpurandare @tianyu-l @weifengpy Summary (human written only) To enable optimizer CUDA graph capture, we need to change learning rate to be a tensor. However, with fsdp2, the optimizer states are DTensor, which cause DTensor sharding propagation materializing the GPU tensors. The root cause is that there is no meta registration for optimizer with learning rate being a tenso…
摘要按规则整理自下方来源原文