* Added SDDMMCOO_hetero support
* removed redundant CUDA kernels
* added benchmark for regression test
* fix
* fixed bug for single src node type
Co-authored-by: Israt Nisa <nisisrat@amazon.com>
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
* [Feature] enable async transfer in NodeDataLoader for homograph
* fix lint issues
* fix device choose when creating stream
* fix test on cpu only machine
* fix pin_memory config
* support homo only
* avoid creating stream in each step and sync via event
* fix lint
* enable graph copy on non-default stream
* fix lint
* refine arg description
* fix conflicts
* Update graph-heterogeneous.rst
`tensor([0, 1, 2, 0, 1, 2])` should be output instead of code
* Update message-api.rst
`updata_all_example()` should be `update_all_example()`
* Update message-efficient.rst
`cat_feat` need to concatenate with `dim=1` for the # edge features to match # edges
* Update nn-construction.rst
all `max_pool` in the aggregator type of `SAGEConv` should be `pool` instead
* Update graph-heterogeneous.rst
`tensor([0, 1, 2, 0, 1, 2])` should be output instead of code
* Update message-api.rst
`updata_all_example()` should be `update_all_example()`
* Update message-efficient.rst
`cat_feat` need to concatenate with `dim=1` for the # edge features to match # edges
* Update nn-construction.rst
all `max_pool` in the aggregator type of `SAGEConv` should be `pool` instead
* Update nn-forward.rst
all `max_pool` in the aggregator type of `SAGEConv` should be `pool` instead
* Update nn-forward.rst
all `max_pool` in the aggregator type of `SAGEConv` should be `pool` instead
Co-authored-by: zhjwy9343 <6593865@qq.com>
* squeeze node labels in FraudDataset
* fix RLModule
* update results in README.md
* fix KeyError in full graph training
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
* add word_ids and simplify
* simplify
* add word_ids to be removed later
* remove word_ids
* seems to work
* tweak
* transpose word_z
* add word_ids example
* check api compatibility
* improve compatibility
* update doc
* tweak verbose
* restore word_z layout; tweak
* tweak
* tweak doc
* word_cT
* use log_weight and some other tweaks
* rewrite README
* update equations
* rewrite for clarity and pass tests
* tweak
* bugfix import
* fix unit test
* fix mult to be the same as old versions
* tweak
* could be a bugfix
* 0/0=nan
* add doc_subgraph utility function
* minor cache optimization
* minor cache tweak
* add environmental variable to trade cache speed for memory
* update README
* tweak
* add sparse update pass unit test
* simplify sparse update
* improve low-memory efficiency
* tweak
* add sample expectation scores to allow resampling
* simplify
* update comment
* avoid edge cases
* bugfix pred scores
* simplify
* add save function
Co-authored-by: Yifei Ma <yifeim@amazon.com>
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
* Based on issue #3436. Improving _SegmentCopyKernel s GPU utilization by switching to nonzero based thread assignment
* fixing lint issues
* Update cub for cuda 11.5 compatibility (#3468)
* fixing type mismatch
* tx guaranteed to be smaller than nnz. Hence removing last check
* minor: updating comment
* adding three unit tests for csr slice method to cover some corner cases
Co-authored-by: Abdurrahman Yasar <ayasar@nvidia.com>
Co-authored-by: nv-dlasalle <63612878+nv-dlasalle@users.noreply.github.com>
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
* relabel gpu
* unittest for ralebl_ on the GPU
* finish Relabel_ for the GPU
* copyright
* re-enable the unittest for edge_subgrah on the GPU
* fix unittest for tensorflow
* use a fixed number of threads
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
Co-authored-by: nv-dlasalle <63612878+nv-dlasalle@users.noreply.github.com>
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
* gpu compact graph template
* cuda compact graph draft
* fix typo
* compact graphs
* pass unit test but fail in training
* example using EdgeDataLoader on the GPU
* refactor cuda_compact_graph and cuda_to_block
* update training scripts
* fix linting
* fix linting
* fix exclude_edges for the GPU
* add --data-cpu & fix copyright
* Add pytorch-direct version
* remove
* add documentation for UnifiedTensor
* Revert "add documentation for UnifiedTensor"
This reverts commit 63ba42644d4aba197c1cb4ea4b85fa1bc43b8849.
* add boundary check for UVM IndexSelect
* relocate boundary check index kernels to cuda
* fix function name
* fix indexkernel in nccl api
* fix argument ordering
* simplify code
* Add a comment for the uvm version
Co-authored-by: shhssdm <shhssdm@gmail.com>
Co-authored-by: xiang song(charlie.song) <classicxsong@gmail.com>
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
* Use ==/!= to compare constant literals (str, bytes, int, float, tuple)
Avoid Syntax Warnings on Python >= 3.8
$ `python3`
```
>>> "" == ""
True
>>> "" is ""
<stdin>:1: SyntaxWarning: "is" with a literal. Did you mean "=="?
True
```
* Use ==/!= to compare constant literals (str, bytes, int, float, tuple)