ciflow/torchtitan/199419: [DCP] Read checkpoint items in parallel in FileSystemReader
- PyTorch: 1769 events in the last 90 days
- PyTorch: 1747th Release in the last 90 days
- Previous: earlier the same day · ciflow/periodic/199127: [UPDATE] > Address review feedback and trim tests
What happened
FileSystemReader reads a rank's items one at a time, and torch.load copies each tensor again. Add thread_count (like FileSystemWriter's): items are read on a thread pool, each decoded from its zip header and directory only, and each tensor's bytes are read straight into the target (through a pinned chunk for accelerator targets). thread_count=1 keeps the serial path; the default splits 32 threads across the node's r…
Summary assembled by rule from the sources below