Da Zheng
|
aaec3d8a0b
|
[Distributed] Support hierarchical partitioning (#3000)
* add.
* fix.
* fix.
* fix.
* fix.
* add tests.
* support node split and edge split.
* support 1 partition.
* add tests.
* fix.
* fix test.
* use hierarchical partition.
* add check.
Co-authored-by: Zheng <dzzhen@3c22fba32af5.ant.amazon.com>
Co-authored-by: Ubuntu <ubuntu@ip-172-31-22-57.us-west-2.compute.internal>
Co-authored-by: Ubuntu <ubuntu@ip-172-31-71-112.ec2.internal>
|
2021-06-16 16:58:23 +08:00 |
|
Chao Ma
|
1aa25cb618
|
update (#2191)
|
2020-09-14 15:13:39 +08:00 |
|
Chao Ma
|
eb9c067b08
|
[Distributed] Copy training scripts in copy_partitions.py (#2010)
* update
* update
* update
* update
* update
* update
* update
* update
* update
* update
* update
* update
* update
* update
|
2020-08-13 12:20:39 +08:00 |
|
Da Zheng
|
4be4b13424
|
[Distributed] add copy_partitions.py (#1866)
* fix bugs.
* eval on both vaidation and testing.
* add script.
* update.
* update launch.
* make train_dist.py independent.
* update readme.
* update readme.
* update readme.
* update readme.
* generate undirected graph.
* rename conf_file to part_config
* use rsync
* make train_dist independent.
Co-authored-by: Ubuntu <ubuntu@ip-172-31-19-1.us-west-2.compute.internal>
Co-authored-by: Ubuntu <ubuntu@ip-172-31-19-115.us-west-2.compute.internal>
Co-authored-by: xiang song(charlie.song) <classicxsong@gmail.com>
|
2020-07-31 11:11:56 -07:00 |
|
Da Zheng
|
02d3197407
|
[Distributed] Pytorch example of distributed GraphSage. (#1495)
* add train_dist.
* Fix sampling example.
* use distributed sampler.
* fix a bug in DistTensor.
* fix distributed training example.
* add graph partition.
* add command
* disable pytorch parallel.
* shutdown correctly.
* load diff graphs.
* add ip_config.txt.
* record timing for each step.
* use ogb
* add profiler.
* fix a bug.
* add train_dist.
* Fix sampling example.
* use distributed sampler.
* fix a bug in DistTensor.
* fix distributed training example.
* add graph partition.
* add command
* disable pytorch parallel.
* shutdown correctly.
* load diff graphs.
* add ip_config.txt.
* record timing for each step.
* use ogb
* add profiler.
* add Ips of the cluster.
* fix exit.
* support multiple clients.
* balance node types and edges.
* move code.
* remove run.sh
* Revert "support multiple clients."
* fix.
* update train_sampling.
* fix.
* fix
* remove run.sh
* update readme.
* update readme.
* use pytorch distributed.
* ensure all trainers run the same number of steps.
* Update README.md
Co-authored-by: Ubuntu <ubuntu@ip-172-31-16-250.us-west-2.compute.internal>
|
2020-06-28 13:09:40 +08:00 |
|