trunk/a2bf3a697dec91d541ceb1816e9ae92355f1f4ad: [cuBLAS] Always eagerly allocate cuBLAS(Lt) workspaces (#194311)
- PyTorch 近 90 天出现 552 次
- PyTorch 近 90 天第 546 次发版
- 上一次:同一天稍早 · trunk/8584bd71f1dbf9b1be692713f4f7c7f322bba923: [torchcomms hash update] update the pinned torchcomms hash (#195793)
发生了什么
authored with codex as discussed w/ @eellison , @ngimel ,~~~ just stashing this prototype here as performance doesn't look great on the hot path:~~~ AI-generated benchmark summary, provided for human review > > Benchmarked cached versus operation-scoped eager cuBLAS workspaces on an NVIDIA GB300, CC 10.3, CUDA 13.4. PyTorch was built for CC 10.0. Measurements are medians across five alternating cached/eager process…
摘要按规则整理自下方来源原文