viable/strict/1791262552: [ROCm] Enable native AsyncTP (#177961)
- PyTorch: 1958 events in the last 90 days
- PyTorch: 1932th Release in the last 90 days
- Previous: earlier the same day · ciflow/trunk/199870
What happened
Summary This PR enables the native AsyncTP path of fused_all_gather_matmul on ROCm (gfx942 / MI300X and gfx950 / MI355X). The native path overlaps the all-gather of the activation shard A with A @ B , using a persistent GEMM that waits on a per-chunk signal before reading each gathered chunk of A . On CUDA this GEMM is a CUTLASS kernel; this PR adds the ROCm equivalent using the Composable Kernel (CK) tile Persisten…
Summary assembled by rule from the sources below