viable/strict/1790725852: [inductor] Partition cross-device fallbacks from CUDA graphs (#190555)
- PyTorch 近 90 天出现 1672 次
- PyTorch 近 90 天第 1653 次发版
- 上一次:同一天稍早 · viable/strict/1790720927: [symm_mem]: Add per-PG stream serialization for ops (#195947)
发生了什么
Summary Split any extern kernel whose FX node reads or produces tensors on more than one device, ignoring meta tensors. This covers FallbackKernel , ExternKernelOut (custom ops with a Tag.out overload), IndexPutFallback , and multi-output ops together with their MultiOutput children. Keep same-device kernels eligible for CUDA graph capture. Rename and simplify the deterministic device_put regression. Add regressions…
摘要按规则整理自下方来源原文