viable/strict/1790890096: [AOTInductor] Parallelize pinned constant staging copies (#195257)
- PyTorch: 1753 events in the last 90 days
- PyTorch: 1731th Release in the last 90 days
- Previous: earlier the same day · trunk/38395f306af10d0f565d5e14f8f55beb3cb44937
What happened
I reviewed the implementation, including the staging-window task division, worker-pool synchronization, fallback paths, and the added CUDA test, and confirmed the controlled performance and correctness results. It is worth noting that this commit introduces no performance regression on the original path. On the 5.39 GB model, after pinning both the loader and the file-backed weight pages to the NUMA node local to th…
Summary assembled by rule from the sources below