viable/strict/1790725852: [inductor] Partition cross-device fallbacks from CUDA graphs (#190555)
- PyTorch: 1666 events in the last 90 days
- PyTorch: 1647th Release in the last 90 days
- Previous: earlier the same day · viable/strict/1790720927: [symm_mem]: Add per-PG stream serialization for ops (#195947)
What happened
Summary Split any extern kernel whose FX node reads or produces tensors on more than one device, ignoring meta tensors. This covers FallbackKernel , ExternKernelOut (custom ops with a Tag.out overload), IndexPutFallback , and multi-output ops together with their MultiOutput children. Keep same-device kernels eligible for CUDA graph capture. Rename and simplify the deterministic device_put regression. Add regressions…
Summary assembled by rule from the sources below