trunk/2208903bd2a09f30786d54b3d2426c98013d0159: [Inductor] Add gfx950 FlyDSL FlexAttention forward (#194309)
- PyTorch: 2130 events in the last 90 days
- PyTorch: 2101th Release in the last 90 days
- Previous: earlier the same day · ciflow/trunk/200436
What happened
I reran these benchmarks on the pinned PyTorch main snapshot and reviewed the results; the updated tables document the final decode coverage and performance. Adds an opt-in FlyDSL FlexAttention forward backend for ROCm gfx950. Summary supports BF16 MHA/GQA prefill and packed-GQA decode supports Q/K dimensions 128 or 192 and V dimension 128 shares K/V staging across query owners, pipelines packed-GQA decode, and spli…
Summary assembled by rule from the sources below