trunk/4a81d5787ea55c5a1cb73afec4cd30a171f8cb5b: [inductor] Gate remaining CUDA TMA tests on CUDA devices (#197141)
- PyTorch: 1070 events in the last 90 days
- PyTorch: 1057th Release in the last 90 days
- Previous: earlier the same day · ciflow/periodic-rocm-mi300/192524: [c10d] Scope the symmetric-memory lifecycle review fixes to ROCm
What happened
Problem When Triton's CPU backend is installed, generic host TMA checks can enable CUDA TMA tests on unsupported GPUs such as A100 and MI300. D119834839 fixed the max-autotune tests, but the same issue remains in AOTInductor and Triton-kernel tests. This Diff Define a shared requires_cuda_tma decorator and apply it to all 11 device-executing TMA tests in AOTInductor and the Triton-kernel suite, while leaving XPU beh…
Summary assembled by rule from the sources below