Rhett Ying
23d575ba81
[Bug] Do not skip graphconv even no edge exists ( #3416 )
2021-11-04 06:14:10 +00:00
Rhett Ying
ee183f69a1
[BugFix] add count_nonzero() into SA_Client ( #3417 )
2021-11-04 06:13:40 +00:00
Rhett Ying
cc83b49e01
[Bug] check dtype before convert to gk ( #3414 )
2021-11-04 06:13:19 +00:00
Rhett Ying
56b04b2125
[BugFix] extract gz into target dir ( #3389 )
2021-11-04 06:11:41 +00:00
Quan (Andy) Gan
84169b1954
[Feature] Graceful handling of exceptions thrown within OpenMP blocks ( #3353 )
...
* graceful c++ exception in OpenMP
* credits
* add test
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com >
2021-11-04 06:10:35 +00:00
Rhett Ying
1c9274f168
[BugFix] initialize data if null when converting from row sorted coo to csr ( #3360 )
2021-11-04 06:08:45 +00:00
Rhett Ying
2b45b8c72d
[Performance] improve coo2csr space complexity when row is not sorted ( #3326 )
...
* [Performance] improve coo2csr space complexity when row is not sorted
* [Perf] replace std::vector<> by NDArray
* keep both impl of unsorted coo to csr and choose according to graph density dynamically
* refine criteria to choose btw Unsorted algos
Co-authored-by: Ubuntu <ubuntu@ip-172-31-34-27.us-west-2.compute.internal >
2021-11-04 06:08:04 +00:00
xiang song(charlie.song)
f9d51fdf0b
[Feature] Add a HINT for the per edge type sampler of heterogeneous DistGraph that highlighting the etypes are sorted already. ( #3260 )
...
* pass cpp test
* distgraph use sorted edge flag.
* lint
* triger
* update test
Co-authored-by: Ubuntu <ubuntu@ip-172-31-2-66.ec2.internal >
2021-11-04 06:02:42 +00:00
esang
f77bee328a
[Bugfix] Fix bugs of farthest_point_sampler ( #3327 )
...
* fix start_idx
* fix the bug when cuda > 0
Co-authored-by: Tong He <hetong007@gmail.com >
2021-11-04 05:48:21 +00:00
Da Zheng
95b1eb734f
[Distributed] Fix a bug in sampling an empty frontier ( #3298 )
...
* handle empty frontiers.
* fix lint.
* fix
Co-authored-by: Ubuntu <ubuntu@ip-172-31-2-202.us-west-1.compute.internal >
Co-authored-by: xiang song(charlie.song) <classicxsong@gmail.com >
2021-11-02 09:51:03 +00:00
xiang song(charlie.song)
4010e20cd9
[Bugfix] Distributed training can not work with dgl.dataloading.negative_sampler ( #3215 )
...
* Fix dist negative data loader bug
* upd
* Fix
Co-authored-by: Da Zheng <zhengda1936@gmail.com >
2021-11-02 09:23:21 +00:00
Da Zheng
f0903275e8
[Distributed] Enable distributed EdgeDataLoader ( #3192 )
...
* make heterogeneous find_edges
* add distributed EdgeDataLoader.
* fix.
* fix a bug.
* fix bugs.
* add tests on distributed heterogeneous graph sampling.
* fix.
Co-authored-by: Zheng <dzzhen@3c22fba32af5.ant.amazon.com >
2021-11-02 09:22:10 +00:00
xiang song(charlie.song)
508197e807
[New Feature] Per edge type sampler for to_homogeneous graphs. ( #3131 )
...
* fix.
* fix.
* fix.
* fix.
* Fix test
* Deprecate old DistEmbedding impl, use synchronized embedding impl
* Basic imple of heterogeneous on homogenenous sampling
* make pass
* Pass C++ test
* Add python test code
* lint
* lint
* Add MultiLayerEtypeNeighborSampler
* Add unitest for single machine dataloader
* Add dist dataloader test for edge type sampler
* Fix lint
* fix
* support for per etype sample
* Fix some bug and enable distributed training with per edge sample
* fix
* Now distributed training works
* turn off some mxnet
* turn off mxnet for some dist test
* fix
* upd
* upd according to the comments
* Fix
* Fix test and now distributed works.
* upd
* upd
* Fix
* Fix bug
* remove dead code.
* upd
* Fix
* upd
* Fix
Co-authored-by: Ubuntu <ubuntu@ip-172-31-71-112.ec2.internal >
Co-authored-by: Ubuntu <ubuntu@ip-172-31-2-66.ec2.internal >
Co-authored-by: Da Zheng <zhengda1936@gmail.com >
2021-11-02 09:21:20 +00:00
Da Zheng
223795eb82
[Distributed] Distributed heterograph training ( #3069 )
...
* support hetero RGCN.
* fix.
* simplify code.
* sample_neighbors return heterograph directly.
* avoid using to_heterogeneous.
* compute canonical etypes in advance.
* fix tests.
* fix.
* fix distributed data loader for heterograph.
* use NodeDataLoader.
* fix bugs in partitioning on heterogeneous graphs.
* fix lint.
* fix tests.
* fix.
* fix.
* fix bugs.
* fix tests.
* fix.
* enable coo for distributed.
* fix.
* fix.
* fix.
* fix.
* fix.
Co-authored-by: Ubuntu <ubuntu@ip-172-31-71-112.ec2.internal >
Co-authored-by: Zheng <dzzhen@3c22fba32af5.ant.amazon.com >
2021-11-02 09:18:15 +00:00
nv-dlasalle
d9a2d74167
[Performance][Feature] Implement edge excluding in EdgeDataLoader on GPU ( #3226 )
...
* Update filter code
* Add unit tests
* Fixes
* Switch to indices
* Rename functions
* Fix linting
* Fix whitespace
* Add doc
* Fix heterograph
* Change workspace allocation
* Fix linting
* Fix docs in filter.py
* Add todo
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com >
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com >
2021-08-20 02:47:48 +00:00
Quan (Andy) Gan
3ccd4a29e8
fix cuda 11.1 crashing bug ( #3265 )
2021-08-20 02:44:56 +00:00
Eric Kim
10beb7b2e0
[Tools] In tools/launch.py, correctly pass all DGL client/server env vars if udf is a multi-command ( #3245 )
...
* Correctly pass all DGL client/server env vars if udf is a multi-command
* Refactor to use wrap_cmd_with_local_envvars() helper fn
2021-08-20 02:42:54 +00:00
Eric Kim
a172b48465
[Tools] Refactor tools/launch.py to handle more python binary names ( #3205 )
...
* Refactors torch dist launcher udf-wrap code to handle more python versions
* minor changes
2021-08-20 02:27:13 +00:00
Rhett Ying
392a2e587d
[bugfix] fix default ntypes/etypes consistency between dgl.DGLGraph and dgl.graph ( #3198 )
2021-08-20 02:18:59 +00:00
xiang song(charlie.song)
d7390763f0
[Distributed] Deprecate old DistEmbedding impl, use synchronized embedding impl ( #3111 )
...
* fix.
* fix.
* fix.
* fix.
* Fix test
* Deprecate old DistEmbedding impl, use synchronized embedding impl
* update doc
Co-authored-by: Ubuntu <ubuntu@ip-172-31-71-112.ec2.internal >
Co-authored-by: Ubuntu <ubuntu@ip-172-31-2-66.ec2.internal >
Co-authored-by: Da Zheng <zhengda1936@gmail.com >
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com >
2021-07-14 00:15:53 +08:00
Mufei Li
f0fafa2062
[Bug fix] Fix batch information with remove_nodes/edges applied ( #3119 )
...
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Fix lint
* Fix
* Update
* Fix test cases
* Fix
* add docstrings
* Update
Co-authored-by: Ubuntu <ubuntu@ip-172-31-28-17.us-west-2.compute.internal >
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com >
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com >
2021-07-13 16:27:31 +08:00
Quan (Andy) Gan
2e19ba8025
add exclude self option for EdgeDataLoader ( #3122 )
...
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com >
2021-07-13 13:30:16 +08:00
Quan (Andy) Gan
b576e617ad
[Feature] Add left normalizer for GCN ( #3114 )
...
* add left normalizer for gcn
* fix
* fixes and some bug stuff
2021-07-13 12:49:35 +08:00
Rhett Ying
186ef59283
[Feature] apply dgl.reorder() onto several node classification datase… ( #3102 )
...
* [Feature] apply dgl.reorder() onto several node classification datasets in DGL
* rebase on latest dgl.reorder_graph()
2021-07-13 08:52:17 +08:00
Rhett-Ying
175f53decf
[Feature] enable edge reorder in dgl.reorder_graph() ( #3113 )
...
* [Feature] enable edge reorder in dgl.reorder_graph()
* refine doc string
* refine doc string for dgl.reorder_graph
* refine doc string further
2021-07-12 14:11:46 +08:00
Rhett-Ying
32d1f3ac97
[BugFix] skip frames whose num_rows is zero in dgl.heterograph.combine_frames() ( #3110 )
...
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com >
2021-07-07 14:15:06 +08:00
Israt Nisa
188152b853
[Feature] Add Heterograph support on Python for builtin unary msg functions (copy_u, copy_e) ( #2989 )
...
* heterograph for binary func
* Added SDDMM support
* Added unittest
* added binary test cases
* unary mfuncs works
* Fixed lint err
* lint check and others
* link check
* fixed import *_hetero issue
* lint check
* replace torch with dgl backend
* lint cehck
* removed torch from test
* skip mxnet unittest
* skip gpu test
* Remove unused/duplicated code
* minor
* changed data structure of ndata and edata
* link check
* reorganized
* minor lint
* minor lint
* raise error for udf func
* lint check
* fix for CUDA 10.1
* add a note for future cross-type max/min reducing
* Add support CUDA < 11
* lint check
* tidied C code
* remove dummy GSDDMM_hetero backward implementation
Co-authored-by: Israt Nisa <nisisrat@amazon.com >
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com >
Co-authored-by: Quan Gan <coin2028@hotmail.com >
2021-07-06 20:41:58 +08:00
Minjie Wang
d3e4460b98
[Doc] Add docstring for missing APIs ( #3088 )
...
* add docstring for missing API; fix some docstring
* rename apis; address comments
2021-07-05 18:40:07 +08:00
Da Zheng
485c04cff6
[Distributed] Fix a few bugs in distributed API ( #3094 )
...
* fix.
* fix.
* fix.
* fix.
* Fix test
Co-authored-by: Ubuntu <ubuntu@ip-172-31-71-112.ec2.internal >
Co-authored-by: Ubuntu <ubuntu@ip-172-31-2-66.ec2.internal >
2021-07-05 15:12:02 +08:00
Da Zheng
0884d02465
[Distributed] Fix bugs in partitioning on heterogeneous graphs. ( #3085 )
...
* fix bugs in partitioning on heterogeneous graphs.
* fix.
* fix.
* fix example.
* fix.
* fix test.
* fix.
* fix.
* fix.
* fix tests.
Co-authored-by: Ubuntu <ubuntu@ip-172-31-71-112.ec2.internal >
Co-authored-by: Zheng <dzzhen@3c22fba32af5.ant.amazon.com >
2021-07-02 21:16:14 +08:00
Jinjing Zhou
0b3a6216f5
[Test] Enable kvstore test ( #3079 )
...
* try enable kvstore test
* fix
* fix
* seperate out kvstore test
* add comment
2021-07-02 14:54:38 +08:00
nv-dlasalle
a0390dde93
[Feature] Add dgl.utils.is_sorted_srcdst() ( #2685 )
...
* Add dgl.utils.is_sorted_srcdst
* Fix linting issues
* delete blank line
* Specify datatype to index tensor in test
* Force integer conversion
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com >
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com >
2021-07-02 11:29:24 +08:00
Rhett-Ying
4307fe8895
[Feature] Add dgl.reorder() to re-order graph according to specified … ( #3063 )
...
* [Feature] Add dgl.reorder() to re-order graph according to specified strategy
* fix unit test failure for metis reorder
* fix unit test failure on mxnet_cpu
* refine unit test for dgl.reorder
* fix unit test failure on mxnet
* fix array_equal error for mxnet unit test
* fix unit test failure for mxnet
* convert metis output to numpy array explicitly
Co-authored-by: Tong He <hetong007@gmail.com >
2021-06-30 13:29:08 +08:00
Jinjing Zhou
9664cdffd3
[Build] Make nccl optional ( #3056 )
...
* fix
* remove nvidiasmi
* fix
* fix docs
* fix
* fix
* 1
* fix
* remove
* skip deprecated kernel
* fix
* Revert "skip deprecated kernel"
This reverts commit c5ceb7f60dbbaf065b81cc3680757fd611d90ad3.
* fix
2021-06-27 22:36:01 +08:00
Quan (Andy) Gan
acd21a6d60
[Feature] Support direct creation from CSR and CSC ( #3045 )
...
* csr and csc creation
* fix
* fix
* fixes to adj transpose
* fine
* raise error if indptr did not match number of nodes
* fix
* huh?
* oh
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com >
2021-06-25 10:05:47 +08:00
xiang song(charlie.song)
2f7ca41459
[Bug fix] Use shared memory for grad sync when NCCL is not avaliable as PyTorch distributed backend. ( #3034 )
...
* Use shared memory for grad sync when NCCL is not avaliable as PyTorch distributed backend.
Fix small bugs and update unitests
* Fix bug
* update test
* update test
* Fix unitest
* Fix unitest
* Fix test
* Fix
* simple update
Co-authored-by: Ubuntu <ubuntu@ip-172-31-24-212.ec2.internal >
2021-06-24 12:06:54 +08:00
Qidong Su
e56bbafd25
[Feature] Biased Neighbor Sampling ( #2987 )
...
* update
* update
* update
* update
* lint
* lint
* update
* update
* update
* update
* update
* update
* update
* update
* update
* update
* update
* update
* update
* update
* lint
* update
* clone
* update
* update
* update
* update
* replace idarray with ndarray
* refactor cpp part
* refactor python part
* debug
* refactor interface
* test and doc
* lint and test
* lint
* fix
* fix
* fix
* const
* doc
* fix
* fix
* fix
* fix
* fix & doc
* fix
* fix
* update
* update
* update
* merge
* doc
* doc
* lint
* fix
* more tests
* doc
* fix
* fix
* update
* update
* update
* fix
* fix
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com >
2021-06-23 17:44:51 +08:00
Quan (Andy) Gan
e667545da5
[Feature] Node2vec ( #2992 )
...
* add seal example
* 1. add paper infomation in examples/README
2. adjust codes
3. option test
* use latest `to_simple` to replace coalesce graph function
* remove outdated codes
* remove useless comment
* Node2vec
1.implement node2vec random walk c++ op
2.implement node2vec model
3.implement node2vec example
* add CMakeLists file modify
* refine c++ codes
* refine c++ codes
* add missing whitespace
* refine python codes
* add codes
* add node2vec_impl.h
* fix codes
* fix code style problem
* fixes
* remove
* lots of changes
* add benchmark
* fixes
Co-authored-by: smilexuhc <smile.xuhc@gmail.com >
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com >
2021-06-23 14:47:56 +08:00
nv-dlasalle
7359481497
[Feature][GPU] Add function for setting weights of a sparse embedding on multiple GPUs. ( #3047 )
...
* add unit test
* Extend NDArrayPartition object
* Add method for setting embedding, and improve documentation
* Sync before returning
* Use name unique to sparse embedding class to avoid delete
Co-authored-by: xiang song(charlie.song) <classicxsong@gmail.com >
2021-06-22 08:48:12 -07:00
Mufei Li
ff519f98c3
[API] Standardize Subgraph APIs ( #2929 )
...
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Fix
* Update
* Fix subgraph tests
* Capture stdout for distributed test
* Capture stdout for distributed test
* Update
* Update
* Update
* Update subgraph.cc
Co-authored-by: Ubuntu <ubuntu@ip-172-31-28-17.us-west-2.compute.internal >
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com >
2021-06-21 19:53:37 +08:00
Kay Liu
9706eaa895
[Feature] add permission information and fix import problems ( #3036 )
...
* [Feature] add positive negative statistics
* [Feature] add permission information and fix import problem
* fix backend incompatible problem
* modify random split to remove sklearn usage
* modify file read to remove pandas usage
* add datasets into doc
* add random seed in data splitting
* add dataset unit test
* usage permission information update
Co-authored-by: zhjwy9343 <6593865@qq.com >
2021-06-21 11:55:15 +08:00
Da Zheng
aaec3d8a0b
[Distributed] Support hierarchical partitioning ( #3000 )
...
* add.
* fix.
* fix.
* fix.
* fix.
* add tests.
* support node split and edge split.
* support 1 partition.
* add tests.
* fix.
* fix test.
* use hierarchical partition.
* add check.
Co-authored-by: Zheng <dzzhen@3c22fba32af5.ant.amazon.com >
Co-authored-by: Ubuntu <ubuntu@ip-172-31-22-57.us-west-2.compute.internal >
Co-authored-by: Ubuntu <ubuntu@ip-172-31-71-112.ec2.internal >
2021-06-16 16:58:23 +08:00
Jinjing Zhou
a303f07819
fix #2952 ( #3010 )
2021-06-14 15:09:14 +08:00
Jinjing Zhou
17141dd372
[Fix] Fix to_homo for graph with zero nodes ntype ( #3011 )
...
* fix #2870
* lint
* fix
2021-06-14 14:34:13 +08:00
Jinjing Zhou
2570d412f9
[NN] Add fast path for GateGCNConv when it has only one edge type ( #2994 )
...
* fix gatedgcn
* fix lint
2021-06-12 01:13:18 +08:00
nv-dlasalle
17d604b5c7
[Feature] Allow using NCCL for communication in dgl.NodeEmbedding and dgl.SparseOptimizer ( #2824 )
...
* Split from NCCL PR
* Fix type in comment
* Expand documentation for sparse_all_to_all_push
* Restore previous behavior in example
* Re-work optimizer to use NCCL based on gradient location
* Allow for running with embedding on CPU but using NCCL for gradient exchange
* Optimize single partition case
* Fix pylint errors
* Add missing include
* fix gradient indexing
* Fix line continuation
* Migrate 'first_step'
* Skip tests without enough GPUs to run NCCL
* Improve empty tensor handling for pytorch 1.5
* Fix indentation
* Allow multiple NCCL communicator to coexist
* Improve handling of empty message
* Update python/dgl/nn/pytorch/sparse_emb.py
Co-authored-by: xiang song(charlie.song) <classicxsong@gmail.com >
* Update python/dgl/nn/pytorch/sparse_emb.py
Co-authored-by: xiang song(charlie.song) <classicxsong@gmail.com >
* Keepy empty tensor dimensionaless
* th.empty -> th.tensor
* Preserve shape for empty non-zero dimension tensors
* Use shared state, when embedding is shared
* Add support for gathering an embedding
* Fix typo
* Fix more typos
* Fix backend call
* Use NodeDataLoader to take advantage of ddp
* Update training script to share memory
* Only squeeze last dimension
* Better handle empty message
* Keep embedding on the target device GPU if dgl_sparse if false in RGCN example
* Fix typo in comment
* Add asserts
* Improve documentation in example
Co-authored-by: xiang song(charlie.song) <classicxsong@gmail.com >
2021-06-10 21:19:00 -07:00
Mufei Li
5be937a7fb
[Kernel] Slicing Batched Graphs ( #2965 )
...
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Add files via upload
* Add files via upload
* Add files via upload
* Add files via upload
* Update
* Update
* Add files via upload
* Add files via upload
* Update
* Lint
* Add files via upload
* Lint
* Update
* Update
* Update
* Update
* Update
* Lint Fix
* Lint
Co-authored-by: Ubuntu <ubuntu@ip-172-31-12-161.us-west-2.compute.internal >
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com >
Co-authored-by: xiang song(charlie.song) <classicxsong@gmail.com >
2021-06-10 12:07:44 +08:00
Tong He
972a9f1323
[Doc] Re-organize the code for dgl.geometry, and expose it in the doc ( #2982 )
...
* reorg and expose dgl.geometry
* fix lint
* fix test
* fix
2021-06-07 11:19:39 +08:00
Da Zheng
b1628f2398
[Tutorial] Distributed node classification. ( #2969 )
...
* add init version.
* fix build.
* fix format.
* fix.
* fix.
* fix format.
* update README.
* avoid running CI on distributed training tutorials.
* Update tutorials/dist/1_node_classification.py
Co-authored-by: xiang song(charlie.song) <classicxsong@gmail.com >
* fix.
Co-authored-by: xiang song(charlie.song) <classicxsong@gmail.com >
2021-06-04 21:50:16 +08:00
Jinjing Zhou
6a56562a7c
[CI] Use k8s cluster ( #2957 )
...
* add
* fix
* set default
* fix
* try master
* try fix
* try
* fix
* 111
* fix
* fix
* update
* ccc
* try
* fix
* fix
* try new machine
* fix
* fix
* fix
* Revert "fix"
This reverts commit e716d66b046f92fe7ae368947a51a036a7a3188a.
* try
* more parrallel
* use k8s for all
* fix name
* try not specify instance type
* ci
* use one yaml
* Revert "use one yaml"
This reverts commit 717d8d852be39fbf2e2e45f9f224aa97907c372c.
* add timeout
* fix permission
* mount efs
* print
* fix pvc
* fix
* restrict num of gpu instances
* check
* fix
* fix
2021-06-04 18:31:11 +08:00