* [Bugfix] fix a compile error for Debug-BuildType on Windows Platform
When using CMakeLists.txt to build the "Debug" BuildType on the Windows Platform, it has three compile errors (C4716) in the file "dgl\src\runtime\shared_mem.cc":
'dgl::runtime::SharedMemory::CreateNew': must return a value
'dgl::runtime::SharedMemory::Open': must return a value
'dgl::runtime::SharedMemory::Exist': must return a value
* [Bugfix] cmake error "cannot find load file" when DGL as a sub_directory on Linux
When using DGL as a subdirectory in a CMake Project, the "CMAKE_SOURCE_DIR" here will return the parent cmake scope dir, which is not a expected dir.
Maybe it is better to use "CMAKE_CURRENT_SOURCE_DIR" to set "GKLIB_PATH".
* [Bugfix] cmd cmake error when DGL as a subdirectory
When DGL as a subdirectory of another project, the WORKING_DIRECTORY of "add_custom_command" will be incorrect at the line 255 of "CMakeLists.txt", such that making a cmake "setlocal" error.
* The start of experiments of Jiahang Li on GraphSAINT.
* a nightly build
* a nightly build
Check the basic pipeline of codes. Next to check the details of samplers , GCN layer (forward propagation) and loss (backward propagation)
* a night build
* Implement GraphSAINT with torch.dataloader
There're still some bugs with sampling in training procedure
* Test validity
Succeed in testing validity on ppi_node experiments without testing other setup.
1. Online sampling on ppi_node experiments performs perfectly.
2. Sampling speed is a bit slow because the operations on [dgl.subgraphs], next step is to improve this part by putting the conversion into parallelism
3. Figuring out why offline+online sampling method performs bad, which does not make sense
4. Doing experiments on other setup
* Implement saint with torch.dataloader
Use torch.dataloader to speed up saint sampling with experiments. Except experiments on too large dataset Amazon, we've done some experiments on other four datasets including ppi, flickr, reddit and yelp. Preliminary experimental results show consumed time and metrics reach not bad level. Next step is to employ more accurate profiler which is the line_profiler to test consumed period, and adjust num_workers to speed up sampling procedures on same certain datasets faster.
* a nightly build
* Update .gitignore
* reorganize codes
Reorganize some codes and comments.
* a nightly build
* Update .gitignore
* fix bugs
Fix bugs about why fully offline sampling and author's version don't work
* reorganize files and codes
Reorganize files and codes then do some experiments to test the performance of offline sampling and online sampling
* do some experiments and update README
* a nightly build
* a nightly build
* Update README.md
* delete unnecessary files
* Update README.md
* a nightly update
1. handle directory named 'graphsaintdata'
2. control graph shift between gpu and cpu related to large dataset ('amazon')
3. remove parameter 'train'
4. refine annotations of the sampler
5. update README.md including updating dataset info, dependencies info, etc
* a nightly update
explain config differences in TEST part
remove a sampling time variant
make 'online' an argument
change 'norm' to 'sampler'
explain parameters in README.md
* Update README.md
* a nightly build
* make online an argument
* refine README.md
* refine codes of `collate_fn` in sampler.py, in training phase only return one subgraph, no need to check if the number of subgraphs larger than 1
* Update sampler.py
check the problem on flickr is about overfitting.
* a nightly update
Fix the overfitting problem of `flickr` dataset. We need to restrict the number of subgraphs (also the number of iterations) used in each epoch of training phase. Or it might overfit when validating at the end of each epoch. The method to limit the number is a formula specified by the author.
* Set up a new flag `full` specifying if the number of subgraphs used in training phase equals to that of pre-sampled subgraphs
* Modify codes and annotations related the new flag
* Add a new parameter called `node_budget` in the base class `SAINTSampler` to compute the specific formula
* set `gpu` as a command line argument
* Update README.md
* Finish the experiments on Flickr, which is done after adding new flag `full`
* a nightly update
* use half of edges in the original graph to do sampling
* test dgl.random.choice with or without replacement with half of edges
~ next is to test what if put the calculating probability part out of __getitem__ can speed up sampling and try to implement sampling method of author
* employ cython to implement edge sampling for per edge
* employ cython to implement edge sampling for per edge
* doing experiments to test consumed time and performance
** the consumed time decreased to approximately 480s, the performance decrease about 5 points.
* deprecate cython implementation
* Revert "employ cython to implement edge sampling for per edge"
* This reverts commit 4ba4f092
* Deprecate cython implementation
* Reserve half-edges mechanism
* a nightly update
* delete unnecessary annotations
Co-authored-by: Mufei Li <mufeili1996@gmail.com>
* [Feature] enable create/set/free cuda stream for internal use
* add unit test
* fix unit test failure on mxnet and tf
* refactor stream wrapper
* fix lint error
* fix lint error
* [Feature] Exclude edges in sample_neighbors
Extending sample_neighbors and sample_frontier
API to support exclude_edges parameter.
exclude_edges support tensor and dict data
Feature enable excluding certain edges
during neighborhood sampling
Exclude_edges contains EID's of edges
which will be excluded
during neighbor picking for seed nodes.
Added test case for heterograph and homograph
RFC issue id: 2944
* compatibility
* fix
* fix
Co-authored-by: Quan Gan <coin2028@hotmail.com>
* [Performance] improve coo2csr space complexity when row is not sorted
* [Perf] replace std::vector<> by NDArray
* keep both impl of unsorted coo to csr and choose according to graph density dynamically
* refine criteria to choose btw Unsorted algos
Co-authored-by: Ubuntu <ubuntu@ip-172-31-34-27.us-west-2.compute.internal>
* publish pct
* add train_cls
* add readme
* update opt for point transformer
* update the example index
* update for comments
Co-authored-by: Tong He <hetong007@gmail.com>
* rgcn with new heterograph API
* apply_edge() forward for multi relation
* undoing changes from rgcn-hetero
* backward apply_edge(copy_u) added
* unittest for apply_edge(copy_e)
* Compatible with new PRs
* resolving conflict with master
* Bringing back change after resolving conflict
* minor
* minor
Co-authored-by: Israt Nisa <nisisrat@amazon.com>
* [CPU, Parallel] Rewriting omp pragmas with parallel_for
* [CPU, Parallel] Decrease number of calls to task function
* c[CPU, Parallel] Modify calls to new interface of parallel_for
* Make column double indices lazy
* Copy indices to proper contexts
* Fix initialization
* Add unit test
* Fix unit test for tensorflow
* Remove unused member
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
* Doc: HetGConv note for `edge_type_subgraph`
For the rational behind this note see the discussion in #3003.
* WIP: Requirements for building the docs.
Had to install all these packages to get further with building the docs.
Didn't successfully build them yet, tho.
* Test
* Update
* Update hetero.py
* Update hetero.py
Co-authored-by: Ubuntu <ubuntu@ip-172-31-13-32.us-west-2.compute.internal>
Co-authored-by: Mufei Li <mufeili1996@gmail.com>
* some modifications for pointnet2
* temporarily save changes
* move files to new directory point_transformer
* implement point transformer for classification
* restore train_cls in pointnet
* implement point transformer for partseg
* fix point transformer for nan loss
* modify point transformer for cls
* modify training setting
* update transformer for cls
* update code
* update code for latest performance
* update the example index
* some minor changes
Co-authored-by: Tong He <hetong007@gmail.com>
* allow for configuring default_dir
* allow for using DGLDEFAULTDIR environment variable
* Update env_var.rst
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>