trunk/bfd8c2497aecccb5e32f08446ac644b87ee2e4d2: [inductor][NVGEMM] Improve NVFP4 autotuning (#194652)
- PyTorch: 1835 events in the last 90 days
- PyTorch: 1811th Release in the last 90 days
- Previous: earlier the same day · trunk/bf06d76b78b5735e759f6f2f4668c4baca964eee: [inductor] Make CUDA graph autotuning measurements consistent (#194649)
What happened
Human commentary: Short NVFP4 inference GEMMs need a small automatic candidate set and measurements that reflect transformer decode execution. AI-assisted content: Problem Fixes #198172 nvMatmulHeuristics does not cover several Blackwell NVFP4 decode winners, while an unrestricted supplemental pool adds substantial compile and profiling cost. The related shape rules were also split between code generation and candid…
Summary assembled by rule from the sources below