提交

提交图

2937 次代码提交

作者 SHA1 备注 提交日期
Xin Yao b717c8bf0d [BugFix] Fix bugs in GPU sampling and enable unit tests for dataloaders on the GPU (#3474)
* enable unit tests for dataloader on the GPU

* fix compatibility

* copyright

* fix linting

Co-authored-by: nv-dlasalle <63612878+nv-dlasalle@users.noreply.github.com>
2021-11-04 10:34:56 -07:00
Xin Yao d3ae7544cd [Feature] aten::Relabel_() for the GPU (#3445)
* relabel gpu

* unittest for ralebl_ on the GPU

* finish Relabel_ for the GPU

* copyright

* re-enable the unittest for edge_subgrah on the GPU

* fix unittest for tensorflow

* use a fixed number of threads

Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
Co-authored-by: nv-dlasalle <63612878+nv-dlasalle@users.noreply.github.com>
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
2021-11-04 08:58:31 -07:00
Mufei Li f46080a4d1 [Feature] k-hop Subgraph Extraction (#3458)
* Update

* Fix

* Fix

* Update

* Update

* Update

* Fix CI

* Fix

* Fix

* Fix

* Update

* Update

* Update

* Fix

* Fix

* Fix for TF
2021-11-04 15:47:35 +08:00
Shaked Brody 64f20eeaec Update CONTRIBUTORS.md (#3475) 2021-11-03 23:05:15 +08:00
Shaked Brody e2f33fd5cc [NN][Model] GATv2 (#3473)
* [Model][Core] GATv2

* lint

* gatv2conv.py

* lint

* lint

* style and docs

* lint

* gatv2conv fix

Co-authored-by: Shaked Brody shakedbr@campus.technion.ac.il <shakedbr@tangerine.cslcs.technion.ac.il>
Co-authored-by: Mufei Li <mufeili1996@gmail.com>
2021-11-03 21:55:56 +08:00
nv-dlasalle f5102145a7 Update cub for cuda 11.5 compatibility (#3468) 2021-11-03 13:14:18 +08:00
Quan (Andy) Gan cac25f636f fix compatibility with PyTorch 1.10 (#3454) 2021-10-29 09:33:08 +08:00
Xin Yao 9067565a25 update contributors (#3451) 2021-10-28 13:19:21 +08:00
Kamil Kamiński 51c6509704 [NN] Add EGATConv nn.module (#3425)
* added nn pytorch egatconv

* aligned with test build

* aligned with test build

* fixed wihite spaces

* fixed wihite spaces

* fixed wihite spaces

* added missing egatconv in imports

* added indentation in forward

* GATConv based implementation

* removed **kw_args

* added dgl relative imports

* PR corrections

* added DGL Error to EGATConv imports

* Update test_nn.py

Co-authored-by: Argusmocny <k.kaminski@cent.uw.edu.pl>
Co-authored-by: Mufei Li <mufeili1996@gmail.com>
2021-10-28 00:38:08 +08:00
Jinjing Zhou a9c83bce15 Fix #3437 (#3440) 2021-10-26 17:56:48 +08:00
Hongyu Cai 579cd3eb49 Update README.md (#3442) 2021-10-26 14:55:09 +08:00
Xin Yao a8c81018c5 [Sampling] Implement dgl.compact_graphs() for the GPU (#3423)
* gpu compact graph template

* cuda compact graph draft

* fix typo

* compact graphs

* pass unit test but fail in training

* example using EdgeDataLoader on the GPU

* refactor cuda_compact_graph and cuda_to_block

* update training scripts

* fix linting

* fix linting

* fix exclude_edges for the GPU

* add --data-cpu & fix copyright
2021-10-20 22:07:35 -07:00
Cheng Wan 308e52a37e [Doc] remove duplicate papers (#3393)
* remove duplicate paper

* Update README.md
2021-10-19 14:41:37 +08:00
nv-dlasalle c560040f5f [Fix] Split nccl sparse push into two groups (#3404) 2021-10-18 18:54:09 +08:00
David Min aa11aaa4ba [Peformance] Parallelize CSRSliceRows() (#3409)
* parallelize CSRRowSlice()

* use parallel_for for the second loop

Co-authored-by: nv-dlasalle <63612878+nv-dlasalle@users.noreply.github.com>
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
2021-10-18 17:25:21 +08:00
Cheng Wan ff94ee80b1 [BugFix] Avoid Memory Leak Issue in PyTorch Backend (#3386)
* try to avoid memory leak

* try to avoid memory leak

* avoid memory leak with no hope

* Revert "avoid memory leak with no hope"

This reverts commit c77befe9479f46758e744642f66dd209b50eef7d.

* no message

* Update sparse.py

* Update tensor.py

Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
2021-10-18 17:24:54 +08:00
HaoWei-TomTom f703941840 [Bugfix][Pytorch] Fix model save and load bug of stgcn_wave (#3303)
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
2021-10-18 16:17:54 +08:00
Quan (Andy) Gan b81bb91465 [Bug] Fix edge exclusion still not working for full neighbor sampling (#3424) 2021-10-15 17:51:56 +08:00
David Min 4f5c3aa234 [Bugfix] Add UVM specialized IndexSelect kernels which perform boundary checks (#3293)
* Add pytorch-direct version

* remove

* add documentation for UnifiedTensor

* Revert "add documentation for UnifiedTensor"

This reverts commit 63ba42644d4aba197c1cb4ea4b85fa1bc43b8849.

* add boundary check for UVM IndexSelect

* relocate boundary check index kernels to cuda

* fix function name

* fix indexkernel in nccl api

* fix argument ordering

* simplify code

* Add a comment for the uvm version

Co-authored-by: shhssdm <shhssdm@gmail.com>
Co-authored-by: xiang song(charlie.song) <classicxsong@gmail.com>
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
2021-10-15 15:05:00 +08:00
Christian Clauss 04ed6126b5 [Fix] Use ==/!= to compare constant literals (str, bytes, int, float, tuple) (#3415)
* Use ==/!= to compare constant literals (str, bytes, int, float, tuple)

Avoid Syntax Warnings on Python >= 3.8

$ `python3`
```
>>> "" == ""
True
>>> "" is ""
<stdin>:1: SyntaxWarning: "is" with a literal. Did you mean "=="?
True
```

* Use ==/!= to compare constant literals (str, bytes, int, float, tuple)
2021-10-14 17:35:37 +08:00
nv-dlasalle b81efb2b2d [PyTorch][Bugfix] Use uint8 instead of bool in pytorch to be compatible with nightly version (#3406)
* Use uint8 instead of bool in pytorch

* Handle type aliases

* Fix syntax error

Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
2021-10-14 17:34:55 +08:00
mszarma a47ab71d44 [DOCS] Add training on CPU sections to docs (#3398) 2021-10-14 14:53:16 +08:00
zexi yuan 188630698e [Bugfix] three bugs related to using DGL as a subdirectory(third_party) of another project. (#3379)
* [Bugfix] fix a compile error for Debug-BuildType on Windows Platform

When using CMakeLists.txt to build the "Debug" BuildType on the Windows Platform, it has three compile errors (C4716) in the file "dgl\src\runtime\shared_mem.cc":

'dgl::runtime::SharedMemory::CreateNew': must return a value
'dgl::runtime::SharedMemory::Open': must return a value
'dgl::runtime::SharedMemory::Exist': must return a value

* [Bugfix] cmake error "cannot find load file" when DGL as a sub_directory on Linux

When using DGL as a subdirectory in a CMake Project, the "CMAKE_SOURCE_DIR" here will return the parent cmake scope dir, which is not a expected dir.
Maybe it is better to use "CMAKE_CURRENT_SOURCE_DIR" to set "GKLIB_PATH".

* [Bugfix] cmd cmake error when DGL as a subdirectory

When DGL as a subdirectory of another project, the WORKING_DIRECTORY of "add_custom_command" will be incorrect at the line 255 of "CMakeLists.txt", such that making a cmake "setlocal" error.
2021-10-14 14:43:23 +08:00
Quan (Andy) Gan 5d4f6bca2a [Fix] Fix edge ID exclusion not working in EdgeDataLoader (#3412) 2021-10-14 14:37:54 +08:00
Rhett Ying 8798872f54 [Bug] Do not skip graphconv even no edge exists (#3416) 2021-10-14 14:34:42 +08:00
Rhett Ying 7c7b60be18 [BugFix] add count_nonzero() into SA_Client (#3417) 2021-10-12 06:17:30 -07:00
Rhett Ying 2d88db5a3c [Bug] check dtype before convert to gk (#3414) 2021-10-12 14:34:40 +08:00
Mufei Li d947287367 [README] Add GNNLens (#3411)
* Update README.md

* Update README.md
2021-10-11 16:27:47 +08:00
Israt Nisa 532eaa879b backward now stores DGLGraph index,not DGLGraph object witattached data (#3410)
Co-authored-by: Israt Nisa <nisisrat@amazon.com>
2021-10-11 14:08:01 +08:00
K aef96dfa34 [Model] Refine GraphSAINT (#3328)
* The start of experiments of Jiahang Li on GraphSAINT.

* a nightly build

* a nightly build

Check the basic pipeline of codes. Next to check the details of samplers , GCN layer (forward propagation) and loss (backward propagation)

* a night build

* Implement GraphSAINT with torch.dataloader

There're still some bugs with sampling in training procedure

* Test validity

Succeed in testing validity on ppi_node experiments without testing other setup.
1. Online sampling on ppi_node experiments performs perfectly.
2. Sampling speed is a bit slow because the operations on [dgl.subgraphs], next step is to improve this part by putting the conversion into parallelism
3. Figuring out why offline+online sampling method performs bad, which does not make sense
4. Doing experiments on other setup

* Implement saint with torch.dataloader

Use torch.dataloader to speed up saint sampling with experiments. Except experiments on too large dataset Amazon, we've done some experiments on other four datasets including ppi, flickr, reddit and yelp. Preliminary experimental results show consumed time and metrics reach not bad level. Next step is to employ more accurate profiler which is the line_profiler to test consumed period, and adjust num_workers to speed up sampling procedures on same certain datasets faster.

* a nightly build

* Update .gitignore

* reorganize codes

Reorganize some codes and comments.

* a nightly build

* Update .gitignore

* fix bugs

Fix bugs about why fully offline sampling and author's version don't work

* reorganize files and codes

Reorganize files and codes then do some experiments to test the performance of offline sampling and online sampling

* do some experiments and update README

* a nightly build

* a nightly build

* Update README.md

* delete unnecessary files

* Update README.md

* a nightly update

1. handle directory named 'graphsaintdata'
2. control graph shift between gpu and cpu related to large dataset ('amazon')
3. remove parameter 'train'
4. refine annotations of the sampler
5. update README.md including updating dataset info, dependencies info, etc

* a nightly update

explain config differences in TEST part
remove a sampling time variant
make 'online' an argument
change 'norm' to 'sampler'
explain parameters in README.md

* Update README.md

* a nightly build

* make online an argument
* refine README.md
* refine codes of `collate_fn` in sampler.py, in training phase only return one subgraph, no need to check if the number of subgraphs larger than 1

* Update sampler.py

check the problem on flickr is about overfitting.

* a nightly update

Fix the overfitting problem of `flickr` dataset. We need to restrict the number of subgraphs (also the number of iterations) used in each epoch of training phase. Or it might overfit when validating at the end of each epoch. The method to limit the number is a formula specified by the author.

* Set up a new flag `full` specifying if the number of subgraphs used in training phase equals to that of pre-sampled subgraphs

* Modify codes and annotations related the new flag

* Add a new parameter called `node_budget` in the base class `SAINTSampler` to compute the specific formula

* set `gpu` as a command line argument

* Update README.md

* Finish the experiments on Flickr, which is done after adding new flag `full`

* a nightly update

* use half of edges in the original graph to do sampling
* test dgl.random.choice with or without replacement with half of edges
~ next is to test what if put the calculating probability part out of __getitem__ can speed up sampling and try to implement sampling method of author

* employ cython to implement edge sampling for per edge

* employ cython to implement edge sampling for per edge
* doing experiments to test consumed time and performance
** the consumed time decreased to approximately 480s, the performance decrease about 5 points.
* deprecate cython implementation

* Revert "employ cython to implement edge sampling for per edge"

* This reverts commit 4ba4f092
* Deprecate cython implementation
* Reserve half-edges mechanism

* a nightly update

* delete unnecessary annotations

Co-authored-by: Mufei Li <mufeili1996@gmail.com>
2021-10-07 19:06:28 +08:00
Rhett Ying f9fd7fd7f7 [BugFix] extract gz into target dir (#3389) 2021-09-30 11:35:40 +08:00
Rhett Ying e234fcfa8f [Feature] enable create/set/free cuda stream for internal use (#3334)
* [Feature] enable create/set/free cuda stream for internal use

* add unit test

* fix unit test failure on mxnet and tf

* refactor stream wrapper

* fix lint error

* fix lint error
2021-09-29 15:35:02 +08:00
Jingcheng Yu 5cf48fc69c [Feature] Implement one thread multiple socket (#3200)
Co-authored-by: JingchengYu94 <jingchengyu94@gmail.com>
2021-09-27 21:45:52 -07:00
xiang song(charlie.song) 179d6aab1e [Distributed] Allow user to pass-in extra env parameters when launching a distributed training task. (#3375)
* Allow user to pass-in extra env parameters when launching a distributed training task.

* Update

* upd

Co-authored-by: xiangsx <xiangsx@ip-10-3-59-214.eu-west-1.compute.internal>
2021-09-23 19:03:33 +08:00
Junwen Yao 367a3a34c4 Fix torch import in example (#3372) 2021-09-23 13:23:05 +08:00
Quan (Andy) Gan a04a8d066e [Feature] Graceful handling of exceptions thrown within OpenMP blocks (#3353)
* graceful c++ exception in OpenMP

* credits

* add test

Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
2021-09-22 13:44:46 +08:00
mszarma bc14829fb3 [Feature] Exclude edges in sample_neighbors (#2971)
* [Feature] Exclude edges in sample_neighbors

Extending sample_neighbors and sample_frontier
API to support exclude_edges parameter.

exclude_edges support tensor and dict data
Feature enable excluding certain edges
during neighborhood sampling
Exclude_edges contains EID's of edges
which will be excluded
during neighbor picking for seed nodes.

Added test case for heterograph and homograph
RFC issue id: 2944

* compatibility

* fix

* fix

Co-authored-by: Quan Gan <coin2028@hotmail.com>
2021-09-22 02:15:56 +08:00
Vikram Sharma ac9261b2a0 [Doc] Added md5sum info for OGB-LSC dataset (#3332)
* Added md5sum for the large dataset files

md5sum helps in validating the correctness of large dataset files once downloaded. 

Refer: https://github.com/snap-stanford/ogb/issues/253
2021-09-21 22:36:53 +08:00
Kay Liu fecd6f3c50 [BugFix] fix typo in fakenews dataset variable name (#3363) 2021-09-21 22:06:26 +08:00
nv-dlasalle 01a2214430 Enable faster validation for pytorch graphsage example (#3361) 2021-09-19 17:25:46 -07:00
jwyyy bc5cac4426 fix broadcast tensor dim (#3351)
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
2021-09-19 21:49:56 +08:00
Rhett Ying bacc9047f2 [BugFix] initialize data if null when converting from row sorted coo to csr (#3360) 2021-09-17 17:07:05 +08:00
nv-dlasalle 2647afc9b3 [Performance][Feature] Add src_nodes paramter to to_block() to avoid cost running unique() when available. (#2973)
* Add lhs_nodes are paremeter to to_block

* Update unit test

* Switch to simplified node conversion

* Switch lhs_nodes to be in/out parameter

* Update docs

Co-authored-by: Da Zheng <zhengda1936@gmail.com>
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
2021-09-16 09:05:16 +08:00
xiang song(charlie.song) ad61a9a550 [Bugfix] And PYTHONPATH in server launch. (#3352)
* put PYTHONPATH in server launch

* remove prints

Co-authored-by: xiangsx <xiangsx@ip-10-3-59-214.eu-west-1.compute.internal>
2021-09-14 17:59:16 +08:00
Rhett Ying f4c79f7f6d [Performance] improve coo2csr space complexity when row is not sorted (#3326)
* [Performance] improve coo2csr space complexity when row is not sorted

* [Perf] replace std::vector<> by NDArray

* keep both impl of unsorted coo to csr and choose according to graph density dynamically

* refine criteria to choose btw Unsorted algos

Co-authored-by: Ubuntu <ubuntu@ip-172-31-34-27.us-west-2.compute.internal>
2021-09-14 10:41:55 +08:00
xiang song(charlie.song) a609b4f023 [Bugfix] Fix #3291 (#3333)
* Fix #3291

* update

* fix

* Unit key

* Fix

Co-authored-by: Ubuntu <ubuntu@ip-172-31-2-66.ec2.internal>
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
2021-09-14 08:57:20 +08:00
sanchit-misra 983a4fdd19 Fixes bug #3312 (#3345)
* Fixes bug #3312

* Fixing lint errors

Co-authored-by: Mufei Li <mufeili1996@gmail.com>
2021-09-13 23:01:35 +08:00
esang 3fef5d27d3 [Model] PCT (#3339)
* publish pct

* add train_cls

* add readme

* update opt for point transformer

* update the example index

* update for comments

Co-authored-by: Tong He <hetong007@gmail.com>
2021-09-13 17:17:24 +08:00
Quan (Andy) Gan e7ea0f5364 Fix openmp header (#3325) 2021-09-13 15:41:28 +08:00
skepsun 26b631805f [Bugfix] Fix Correct&Smooth (#3329)
* Update model.py

fix typo

* Update main.py

fix autoscale

* Update README.md

Co-authored-by: Mufei Li <mufeili1996@gmail.com>
2021-09-13 11:46:14 +08:00