← Back to events
ActiveAIRelease

viable/strict/1791262552: [ROCm] Enable native AsyncTP (#177961)

What happened

Summary This PR enables the native AsyncTP path of fused_all_gather_matmul on ROCm (gfx942 / MI300X and gfx950 / MI355X). The native path overlaps the all-gather of the activation shard A with A @ B , using a persistent GEMM that waits on a per-chunk signal before reading each gathered chunk of A . On CUDA this GEMM is a CUTLASS kernel; this PR adds the ROCm equivalent using the Composable Kernel (CK) tile Persisten…

Summary assembled by rule from the sources below

Why it's spreading

Sources