viable/strict/1790977859: [ROCm] Enable test_grouped_mm on ROCm (#199048)
- PyTorch: 1839 events in the last 90 days
- PyTorch: 1815th Release in the last 90 days
- Previous: earlier the same day · viable/strict/1790976466
What happened
DistMatrixOpsTest.test_grouped_mm was skipped on ROCm via @unittest.skipIf(TEST_WITH_ROCM, "ROCm doesn't support CUTLASS"). That reason is a misnomer for this test: it is pure bf16, and on ROCm the backend="cutlass" parametrization routes to the arch-independent _grouped_mm fallback (CUTLASS is never invoked), while the backend="cublaslt" variants self-skip via an existing CUDA-Toolkit-version guard. Remove the ROCm…
Summary assembled by rule from the sources below