viable/strict/1790890096: [AOTInductor] Parallelize pinned constant staging copies (#195257)
- PyTorch 近 90 天出现 1765 次
- PyTorch 近 90 天第 1743 次发版
- 上一次:同一天稍早 · trunk/38395f306af10d0f565d5e14f8f55beb3cb44937
发生了什么
I reviewed the implementation, including the staging-window task division, worker-pool synchronization, fallback paths, and the added CUDA test, and confirmed the controlled performance and correctness results. It is worth noting that this commit introduces no performance regression on the original path. On the 5.39 GB model, after pinning both the loader and the file-backed weight pages to the NUMA node local to th…
摘要按规则整理自下方来源原文