trunk/bfd8c2497aecccb5e32f08446ac644b87ee2e4d2: [inductor][NVGEMM] Improve NVFP4 autotuning (#194652)
- PyTorch 近 90 天出现 1839 次
- PyTorch 近 90 天第 1815 次发版
- 上一次:同一天稍早 · trunk/bf06d76b78b5735e759f6f2f4668c4baca964eee: [inductor] Make CUDA graph autotuning measurements consistent (#194649)
发生了什么
Human commentary: Short NVFP4 inference GEMMs need a small automatic candidate set and measurements that reflect transformer decode execution. AI-assisted content: Problem Fixes #198172 nvMatmulHeuristics does not cover several Blackwell NVFP4 decode winners, while an unrestricted supplemental pool adds substantial compile and profiling cost. The related shape rules were also split between code generation and candid…
摘要按规则整理自下方来源原文