viable/strict/1791287563: [inductor-perf] allow fusion into template epilogue inputs. (#198809)
- PyTorch: 1965 events in the last 90 days
- PyTorch: 1939th Release in the last 90 days
- Previous: earlier the same day · viable/strict/1791284024
What happened
Template prologue fusion lets a pointwise producer be fused into the code where a template loads an input (load_input()). Inputs that a template consumes only in its epilogue, like the addmm/baddbmm bias, could not take a fused producer. A computed bias was always written to memory by a separate kernel and read back in the template epilogue: tmp = bias * 2.0 - 1.0 addmm(tmp, a, b) This PR lets producers of those inp…
Summary assembled by rule from the sources below