viable/strict/1787655491: [Inductor] Add gfx950 FlyDSL GEMM and torch.mm autotuning (#190903)
- PyTorch: 45 events in the last 90 days
- PyTorch: 45th Release in the last 90 days
- Previous: earlier the same day · trunk/94055e9ef64834496cd136a1bf5ddd2621a5be83: [distributed] Run legacy-only NCCL tests with nccl-legacy (#193240)
What happened
Build on the FlyDSL template infrastructure to integrate an FP16/BF16 gfx950 GEMM into torch.mm max-autotune. Summary vendor the gfx950 FlyDSL GEMM kernel and generated template wrapper add default and exhaustive tile-configuration heuristics preserve runtime row strides and storage offsets for supported NT inputs filter shape-incompatible configurations before autotuning reject inputs whose origins or row strides v…
Summary assembled by rule from the sources below