pytorch/torch/testing/_internal/distributed
Rohan Varma b22abbe381 Enable test_distributed to work with spawn mode (#41769)
Summary:
Pull Request resolved: https://github.com/pytorch/pytorch/pull/41769

Currently the tests in `test_distributed` only work with the `fork` mode multiprocessing, this PR introduces support for `spawn` mode multiprocessing as well (while keeping the `fork` mode intact).

Motivations for the change:
1) Spawn multiprocessing is the default on MacOS, so it better emulates how MacOS users would use distributed
2) With python 3.8+, spawn is the default on linux, so we should have test coverage for this
3) PT multiprocessing suggests using spawn/forkserver over fork, for sharing cuda tensors: https://pytorch.org/docs/stable/multiprocessing.html
4) Spawn is better supported with respect to certain sanitizers such as TSAN, so adding this sanitizer coverage may help us uncover issues.

How it is done:
1) Move `test_distributed` tests in `_DistTestBase` class to a shared file `distributed_test` (similar to how the RPC tests are structured)
2) For `Barrier`, refactor the setup of temp directories, as the current version did not work with spawn, each process would get a different randomly generated directory and thus would write to different barriers.
3) Add all the relevant builds to run internally and in OSS.
Running test_distributed with spawn mode in OSS can be done with:
`python test/run_test.py -i distributed/test_distributed_spawn -v`

Reviewed By: izdeby

Differential Revision: D22408023

fbshipit-source-id: e206be16961fd80438f995e221f18139d7e6d2a9
2020-09-08 23:11:12 -07:00
..
nn Explicitly forbidden the other inherited methods of RemoteModule. (#43895) 2020-09-05 14:48:56 -07:00
rpc [RPC profiling] Add test to ensure using record_function works for RPC (#43657) 2020-08-31 11:43:09 -07:00
__init__.py remediation of S205607 2020-07-17 17:19:47 -07:00
ddp_under_dist_autograd_test.py
distributed_test.py Enable test_distributed to work with spawn mode (#41769) 2020-09-08 23:11:12 -07:00
rpc_utils.py [RPC tests] Run DdpUnderDistAutogradTest and DdpComparisonTest with fork too (#42528) 2020-08-05 15:10:29 -07:00