Rhett Ying
4dd16f5dd8
[BugFix] enable DistGraph.find_edge() works with str or tuple of str ( #4319 )
2022-08-01 20:47:15 +08:00
Xin Yao
44b6864114
[Feature] Enable UVA for Weighted Samplers ( #4314 )
...
* enable use for weighted neighbor sampler and biased random walk
* add unit tests
* fix for mxnet/tf
* fix typo
2022-08-01 17:56:51 +08:00
Rhett Ying
d6957c28ff
[BugFix] fix incorrect _bias and bias usage ( #4310 )
2022-07-30 17:40:54 +08:00
Rhett Ying
b3242e90b8
[CI] separate distributed tests from torch cpu tests ( #4313 )
...
* [CI] separate distributed tests from torch cpu tests
* remove TF related env
2022-07-30 16:32:35 +08:00
Xin Yao
86c81b4e92
[Feature] Add CUDA Weighted Neighborhood Sampling ( #4064 )
...
* add weighted sampling without replacement (A-Chao)
* improve Algorithm A-Chao with block-wise prefix sum
* correctly fill out_idxs
* implement weighted sampling with replacement
* small fix
* merge host-side code of weighted/uniform sampling
* enable unit tests for cuda weighted sampling
* move thrust/cub wrapper to the cmake file
* update docs accordingly
* fix linting
* fix linting
* fix unit test
* Bump external CUB/Thrust versions
* Fix code style and update description of algorithm design
* [Feature] GPU support weighted graph neighbor sampling
commit by pengqirong(OPPO)
* merge pengqirong's implementation
* revert the change to cub and thrust
* fix linting
* use DeviceSegmentedSort for better performance
* add more comments
* add necessary notes
* add necessary notes
* resolve some comments
* define THRUST_CUB_WRAPPED_NAMESPACE
* fix doc
Co-authored-by: 彭齐荣 <657017034@qq.com >
2022-07-29 11:08:48 +08:00
Rhett Ying
17f1432ab2
[DistTest] fix incorrect shell if statement ( #4304 )
...
* [DistTest] fix incorrect shell if statement
* fix incorrect use of dist.initialize()
2022-07-28 20:14:47 +08:00
Pengfei Xia
2cf05c5342
[Transform] Allow add data to self loop created by AddSelfLoop or add_self_loop ( #4261 )
...
* Update
* Update functional.py
* Update
* Update test_transform.py
* Update
* Update functional.py
* Update functional.py
* Update functional.py
* Update functional.py
* Update
* Update
* Update functional.py
* Update functional.py
* Update functional.py
* Update functional.py
* Update module.py
* Update test_transform.py
* Update test_transform.py
Co-authored-by: Mufei Li <mufeili1996@gmail.com >
2022-07-27 15:14:03 +08:00
Rhett Ying
069068aa7d
[Log] fix confusing error log in TCPSocket::Bind() ( #4299 )
...
* [Log] fix confusing error log in TCPSocket::Bind()
* fix lint
2022-07-27 08:25:42 +08:00
Dewvin
7e6a6b4ac0
[Feature] Add CUDA Weighted Randomwalk Sampling ( #4243 )
...
* [Feature] Add CUDA Weighted Randomwalk Sampling
* [Feature] Add CUDA Weighted Randomwalk Sampling
* [Feature] Add CUDA Weighted Randomwalk Sampling
* [Feature] Add CUDA Weighted Randomwalk Sampling
* fix empty prob array && enable non-uniform for restart && enable unit tests
* update doc and guide for randomwalk and pinsage
* update comments
Co-authored-by: zhenliangqiu <ubuntu@ip-172-31-24-245.ap-southeast-1.compute.internal >
Co-authored-by: xiny <xiny@nvidia.com >
2022-07-26 13:42:39 +08:00
Serge Panev
5fc1d0c83a
[Dist][Test] Add tests for multi-node DistEmbedding ( #4256 )
...
Signed-off-by: Serge Panev <spanev@nvidia.com >
Co-authored-by: Rhett Ying <85214957+Rhett-Ying@users.noreply.github.com >
2022-07-20 20:04:26 +08:00
Serge Panev
4dc5728af3
[Dist][Test] Improves DistTensor test for num_part id > 2 ( #4265 )
...
Signed-off-by: Serge Panev <spanev@nvidia.com >
Co-authored-by: Rhett Ying <85214957+Rhett-Ying@users.noreply.github.com >
2022-07-19 17:53:41 +08:00
Xin Yao
79b0a50afc
[Unittest][Fix] Several unit tests fixes for Ampere+ and PyTorch 1.12+ ( #4213 )
...
* Fix test_csrmm for tensor core
* unset allow tf32 flag
* update test unified tensor
* skip fp16 for CPU
2022-07-14 19:59:57 +08:00
Rhett Ying
14e4e1b0f7
[Dist] enable to specify sort_etype for sample_etype_neighbours ( #4212 )
...
* [Dist] enable to specify sort_etype for sample_etype_neighbours
* fix lint
* pass argument instead of env
* fix lint and doc string
* refine args
* remove unnecessary lines
* debug only
* debug add sort time log
* change interface
* fix typo
Co-authored-by: Xin Yao <xiny@nvidia.com >
2022-07-11 18:37:25 +08:00
Rhett Ying
c65d6fa55f
[Dist] format dtypes when loading graph in server ( #4228 )
...
* [Dist] format dtypes when loading graph in server
* add test
* refine
* add comments
2022-07-11 14:08:17 +08:00
Serge Panev
e3b6ac8e6c
[Dist][Test] Add tests for multi-node DistTensor ( #4122 )
...
Signed-off-by: Serge Panev <spanev@nvidia.com >
Co-authored-by: Rhett Ying <85214957+Rhett-Ying@users.noreply.github.com >
2022-07-08 10:47:41 +08:00
Vibhu Jawa
9ae117d3cd
[Feature] Add to_cugraph and from_cugraph to DGL ( #4166 )
...
* Added to_cugraph and from_cugraph functionality
* fix from_cugraph example
* Addressed Reviews
* Fix linting
* Apply suggestions to docstrings from code review
Co-authored-by: Minjie Wang <minjie.wang@nyu.edu >
* Add API docs and remove `from_cugraph` alias
* move cugraph tests to tests/cugraph/test_basics.py
* remove ununsed imports from test_basics.py
* remove pytest.importorskip as no longer needed
Co-authored-by: Minjie Wang <minjie.wang@nyu.edu >
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com >
2022-07-07 11:52:17 +08:00
Chang Liu
885be1784b
[Example][Refactor] Refactor GCN example ( #4160 )
...
* Refactor GCN example
* Refactor GCN based on graphsage
* Readme update
* Minor update
* update
* Remove user-defined GCN implementation
* README update
* Update
* Update CONTRIBUTORS.md
* update task_example_test
Co-authored-by: Xin Yao <xiny@nvidia.com >
2022-07-05 15:58:31 +08:00
Rhett Ying
85f281170f
[CI] Add new CI stage for testing cugraph ( #4171 )
...
* [CI] add new stage specific forcuda related features based on nvidia+pytorch
* build and test for gpu_nv
* fix build failure
* fix unit tests
* make -j
* install cython beforehand
* copy cython lib
* test cugraph tests only
* fix typo
* separate test script for cugraph
* refactor build dgl shell
2022-07-05 10:57:07 +08:00
Rhett Ying
dcf1699280
[BugFix] check whether etype sorted when sampling ( #4198 )
2022-07-01 15:52:14 +08:00
Rhett Ying
6a6597a02a
[Feature] extend sort_csr/csc_by_tag to edge ( #4164 )
...
* [Feature] extend sort_csr/csc_by_tag to edge
* fix test ffailure in tensorflow
* refine sorting by edges
* fix docstring
* remove unnecessary mem
Co-authored-by: Xin Yao <xiny@nvidia.com >
2022-07-01 13:16:11 +08:00
Minjie Wang
f7dae4533a
[CI] Reduce CI workload ( #4196 )
...
* try optimize CI
* fix go test; adjust timing report
* disable certain tests for mx/tf backends
* fix ut
* add pydantic
2022-06-30 20:01:16 +08:00
nv-dlasalle
d2a22984c1
[bugfix] Implement __setstate__ for Column ( fixes #4107 ) ( #4174 )
...
* * Workaround for graph data saving/loading compatibility problem in Column class. There may be more places in DGL with the same issue, due to using Python serialization, instead of a more cohesive, comprehensive strategy. This is just a local fix.
* Add checking for non-empty states
* Add unit test
* Handle the case of columns without storage
Co-authored-by: ndickson <ndickson@nvidia.com >
Co-authored-by: Xin Yao <xiny@nvidia.com >
2022-06-30 10:10:44 +08:00
Mufei Li
8b19c28788
Update test_transform.py ( #4190 )
2022-06-29 21:03:23 +08:00
nv-dlasalle
1dddaad4f0
[bugfix] Allow communicators of size one when NCCL is missing ( #3713 )
...
* Update nccl communicator for when NCCL is missing
* Use static_cast
* Add doc string
* Fix whitespace
* Resrtict unit test to GPU runs
Co-authored-by: Xin Yao <xiny@nvidia.com >
2022-06-29 11:48:16 +08:00
Mufei Li
a25a14f2fa
[Bug Fix] Fix A Bug Related to GroupRevRes ( #4181 )
...
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com >
Co-authored-by: Xin Yao <xiny@nvidia.com >
2022-06-28 21:17:20 +08:00
Mufei Li
cb39dbfa13
Update ( #4178 )
2022-06-28 15:44:03 +08:00
nv-dlasalle
020f02498c
[Performance][Optimizer] Enable using UVA and FP16 with SparseAdam Optimizer ( #3885 )
...
* Add uva by default to embedding
* More updates
* Update optimizer
* Add new uva functions
* Expose new pinned memory function
* Add unit tests
* Update formatting
* Fix unit test
* Handle auto UVA case when training is on CPU
* Allow per-embedding decisions for whether to use UVA
* Address spares_optim.py comments
* Remove unused templates
* Update unit test
* Use dgl allocate memory for pinning
* allow automatically unpin
* workaround for d2h copy with a different dtype
* fix linting
* update error message
* update copyright
Co-authored-by: Xin Yao <xiny@nvidia.com >
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com >
2022-06-24 09:20:08 +08:00
Xin Yao
077e002fe5
[Bugfix][Rework] Automatically unpin tensors pinned by DGL (rework #3997 ) ( #4135 )
...
* Explicitly unpin tensoradapter allocated arrays
* Undo unrelated change
* Add unit test
* update unit test
* add pinned_by_dgl flag to NDArray::Container
* use dgl.ndarray for holding the pinning status
* update multi-gpu uva inference
* reinterpret cast NDArray::Container* to DLTensor* in MoveAsDLTensor
* update unpin column and examples
* add unit test for unpin column
Co-authored-by: Dominique LaSalle <dlasalle@nvidia.com >
Co-authored-by: nv-dlasalle <63612878+nv-dlasalle@users.noreply.github.com >
2022-06-23 13:56:54 +08:00
Quan (Andy) Gan
71157b05a8
[Bug] Fix problem with ShaDowKHopSampler working with reverse edge type exclusion ( #4145 )
...
* fix
* fix
* Update utils.py
2022-06-22 21:37:13 +08:00
Mufei Li
31e4a89b23
[DGL-Go] Inference for Node Prediction Pipeline (full & ns) ( #4095 )
...
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
* Update
2022-06-21 17:45:44 +08:00
Rhett Ying
69226588a5
[Dist] defer to load node/edge feats ( #4143 )
...
* [Dist] defer to load node/edge feats
* fix lint
* Update python/dgl/distributed/partition.py
Co-authored-by: Minjie Wang <minjie.wang@nyu.edu >
* Update python/dgl/distributed/partition.py
Co-authored-by: Minjie Wang <minjie.wang@nyu.edu >
* fix lint
Co-authored-by: Minjie Wang <minjie.wang@nyu.edu >
2022-06-20 19:44:37 +08:00
Rhett Ying
b258729b3f
[Dist] set socket as default backend for RPC ( #4120 )
...
* [Dist] set socket as default backend for RPC
* add tests both for socket and tensorpipe
2022-06-16 13:08:19 +08:00
Serge Panev
652f4c0743
[Dist] Add env var for non-default SSH configs in tests ( #4098 )
...
Signed-off-by: Serge Panev <spanev@nvidia.com >
Co-authored-by: Rhett Ying <85214957+Rhett-Ying@users.noreply.github.com >
2022-06-15 12:38:50 +08:00
RecLusIve-F
defa292bc0
[Dataset] Add Flickr and Yelp dataset ( #4099 )
...
* Add Flickr and Yelp dataset
* Update flickr.py
* update
* Update yelp.py
* Update yelp.py
* update
* Update yelp.py
* Update test_data.py
* Update yelp.py
* update
* Update test_data.py
* Update yelp.py
Co-authored-by: Mufei Li <mufeili1996@gmail.com >
2022-06-14 16:14:12 +08:00
Rhett Ying
9501ed6a07
[Dist] master port should be fixed for all trainers ( #4108 )
...
* [Dist] master port should be fixed for all trainers
* add tests for tools/launch.py
2022-06-14 11:22:22 +08:00
Huarui HE
92e7733065
[dataset] Add a reorder flag to builtin datasets ( #4104 )
...
* add argument reorder=False for citation_graph
* add description of the argument reorder
* add reordered/un_reordered save_path
* add version number postfix
Co-authored-by: Mufei Li <mufeili1996@gmail.com >
2022-06-12 12:35:30 +08:00
Rhett Ying
abcc9cce83
disable multiple groups tests due to random failure in CI ( #4101 )
2022-06-09 17:40:53 +08:00
Rhett Ying
cac3720b48
[Dist] enable time out when fetching msg ( #4043 )
...
* [ist] enable time out when fetching msg
* fix lint error
* minor refinements
* improve minor log
* fix dist test
* fix timeout issue in tensorpipe
2022-06-08 20:20:03 +08:00
Rhett Ying
2de80dde7d
[DistTest] add python test of RPC ( #4093 )
...
* [DistTest] add python test of RPC
* remove return
2022-06-08 17:03:02 +08:00
Rhett Ying
c1ff4c9b41
[DistTest] add basic pipeline for dist test across machines ( #3984 )
...
* [DistTest] add basic pipeline for dist test across machines
* move launch remote cmd to separate file
* add test for rpc
* fix function naming rule
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com >
2022-06-08 12:37:02 +08:00
ndickson-nvidia
eabcc58e41
[Bug][Feature] Added cublasGemm<__half> specialization ( #3988 ) ( #4029 )
...
* * Added specialization of cublasGemm function for `__half` type, to try to address https://github.com/dmlc/dgl/issues/3988
* * Added USE_FP16 guard
* * Added test cases to test_segment_mm, to test newly-added FP16 specialization of cublasGemm
* * Replaced for loop in test_segment_mm with pytest.mark.parametrize, as recommended
Co-authored-by: Xin Yao <xiny@nvidia.com >
2022-06-07 16:00:02 +08:00
Mufei Li
6a91d18151
[DGL-Go] Graph Property Prediction Pipeline ( #3927 )
...
* Update
* Update
* Update
* Fix
* Update
* Update
* Update
* Update
* Update
* Fix
* Update
* Update
* Update
* Fix
* Update
* Update
* update
* Update
* Update
* Update
* Fix
* Fix
* Update
* Update
* Update
* Update
* lr_scheduler
* Update
* Update
* Update
* Update
* Update
* Update
* Fix
* Fix
* Fix
* Fix
* Update
* Update
* Update
* Update
* Update
* Fix
* Update
* Update
* Update
* Update
* Fix
* Update
* Update
* Update
* Fix
* Fix
* Fix
* Update
* Update
* Fix
* Fix
* Update
* Update
* Update
* Update
* Update
* Update
* update
* CI
* Update
* Update
* Update
* Update
* Update
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com >
2022-06-04 15:43:34 +08:00
Riju Mukherjee
efd909e62e
[NN] Enhance EGATConv branch ( #4062 )
...
* enhance EGATConv| nfeats as tuples
* egatconv modified for bipartite graphs
* modified docstrings
* added/modified unittests for EGATConv
* Update egatconv.py
* rectified lint errors
Co-authored-by: rijulizer <riju.mukherjee@gmail.com >
Co-authored-by: Mufei Li <mufeili1996@gmail.com >
2022-06-03 11:53:10 +08:00
RecLusIve-F
89655cfda2
[Dataset] Add WikiCS Dataset ( #4035 )
...
* Fix bugs & Update dataset
* Update
* Update wikics.py
* Update wikics.py
* Update test_data.py
* Update wikics.py
* Update wikics.py
* Update wikics.py
* update
* Update module.py
* Update dgl.data.rst
* Update wikics.py
* Update wikics.py
Co-authored-by: Mufei Li <mufeili1996@gmail.com >
2022-06-02 23:03:17 +08:00
Mufei Li
d9c25521bc
[Data] AsGraphPredDataset ( #4073 )
...
* Update
* CI
* Update
* Update
* Fix
* Fix
2022-06-02 16:18:48 +08:00
Quan (Andy) Gan
00c09b9f91
Revert "[bugfix] Explicitly unpin tensoradapter allocated arrays ( #3997 )" ( #4061 )
...
This reverts commit fdd1fe1908 .
2022-05-28 20:44:20 +08:00
Quan (Andy) Gan
c577dc9fc3
add sanity check ( #4050 )
2022-05-28 18:51:33 +08:00
nv-dlasalle
7a065a9c56
[Build][Tests] Enable FP16 for GPU builds in CI ( #4030 )
...
* Enable FP16 for GPU builds in CI
* Limit default GPU archs to pascal and above
* Disable FP16 dispatching for cuda architectures less than 60
* Fix linting
* Fix typos
2022-05-26 09:48:28 -07:00
Minjie Wang
3c129ad71f
[Bugfix] Cython CAPI holding GIL causes deadlock when Python callback is asynchronous ( #4036 )
...
* cython nogil
* move APIs to internal and add unit test
* fix lint
* disable callback array test
2022-05-25 10:02:44 +08:00
Mufei Li
230b886ec5
[Bug fix] Misc Fix for Transforms and NN Modules ( #4038 )
...
* Update module.py
* Update utils.py
* Update utils.py
* Update utils.py
* Update module.py
* Update
* Update
* Update
2022-05-24 18:08:08 +08:00