trunk/d25057513dc155ee32bfdf4e760da11652269423: [inductor][NVGEMM] Fuse NVFP4 output scaling (#194651)
- PyTorch 近 90 天出现 1839 次
- PyTorch 近 90 天第 1815 次发版
- 上一次:同一天稍早 · trunk/bfd8c2497aecccb5e32f08446ac644b87ee2e4d2: [inductor][NVGEMM] Improve NVFP4 autotuning (#194652)
发生了什么
Human commentary: ModelOpt NVFP4 linear layers multiply _scaled_mm output by a scalar dequantization factor. Folding that factor before QKV fan-out removes one standalone scale launch per layer and lets autotuning consider the GEMM implementation that applies it directly. AI-assisted content: Problem Fixes #198173 ModelOpt NVFP4 linears materialize a separate pointwise multiply after _scaled_mm . That launch is espe…
摘要按规则整理自下方来源原文