* Update graphsage multi-gpu example to use mutliple GPUs for validation and
testing.
* Remove argmax
* Fix rebase error
* Add more documentation to example and simplify
* Switch to name shared memory
* Add comment about how training is distributed
* Restore iteration count
* fix munmap error reporting for better error messages
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
* huuuuge update
* remove
* lint
* lint
* fix
* what happened to nccl
* update multi-gpu unsupervised graphsage example
* replace most of the dgl.mp.process with torch.mp.spawn
* update if condition for use_uva case
* update user guide
* address comments
* incorporating suggestions from @jermainewang
* oops
* fix tutorial to pass CI
* oops
* fix again
Co-authored-by: Xin Yao <xiny@nvidia.com>
* WIP: TypedLinear and new RelGraphConv
* wip
* further simplify RGCN
* a bunch of tweak for performance; add basic cpu support
* update on segmm
* wip: segment.cu
* new backward kernel works
* fix a bunch of bugs in kernel; leave idx_a for future
* add nn test for typed_linear
* rgcn nn test
* bugfix in corner case; update RGCN README
* doc
* fix cpp lint
* fix lint
* fix ut
* wip: hgtconv; presorted flag for rgcn
* hgt code and ut; WIP: some fix on reorder graph
* better typed linear init
* fix ut
* fix lint; add docstring
* Add RNAGlib to examples and DGL-powered-projects
* Add RNAGlib to examples and DGL-powered-projects
* Add RNAGlib to examples and DGL-powered-projects
* implement pin_memory/unpin_memory/is_pinned for dgl.graph
* update python docstring
* update c++ docstring
* add test
* fix the broken UnifiedTensor
* XPU_SWITCH for kDLCPUPinned
* a rough version ready for testing
* eliminate extra context parameter for pin/unpin
* update train_sampling
* fix linting
* fix typo
* multi-gpu uva sampling case
* disable new format materialization for pinned graphs
* update python doc for pin_memory_
* fix unit test
* UVA sampling for link prediction
* dispatch most csr ops
* update graphsage example to combine uva sampling and UnifiedTensor
* update graphsage example to combine uva sampling and UnifiedTensor
* update graphsage example to combine uva sampling and UnifiedTensor
* update doc
* update examples
* change unitgraph and heterograph's PinMemory to in-place
* update examples for multi-gpu uva sampling
* update doc
* fix linting
* fix cpu build
* fix is_pinned for DistGraph
* fix is_pinned for DistGraph
* update graphsage unsupervised example
* update doc for gpu sampling
* update some check for sampling device switching
* fix linting
* adapt for new dataloader
* fix linting
* fix
* fix some name issue
* adjust device check
* add unit test for uva sampling & fix some zero_copy bug
* fix linting
* update num_threads in graphsage examples
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
* Adding initial files of example
* Removing old timing code
* Improving doc strings and fixing some minor bugs
* Merging from upstream and addressing PR comments
Co-authored-by: zakjost <jostza@amazon.com>
Co-authored-by: Mufei Li <mufeili1996@gmail.com>
* added distgnn plus libra codebase
* Dist application codes
* added comments in partition code. changed the interface of partitioning call.
* updated readme
* create libra partitioning branch for the PR
* removed disgnn files for first PR
* updated kernel.cc
* added libra_partition.cc and moved libra code from kernel.cc to libra_partition.cc
* fixed lint error; merged libra2dgl.py and main_Libra.py to libra_partition.py; added graphsage/distgnn folder and partition script.
* removed libra2dgl.py
* fixed the lint error and cleaned the code.
* revisions due to PR comments. added distgnn/tools contains partitions routines
* update 2 PR revision I
* fixed errors; also improved the runtime by 10x.
* fixed minor lint error
* fixed some more lints
* PR revision II changed the interface of libra partition function
* rewrite docstring
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
* rgcn with new heterograph API
* added new apply_edge()
* optimized forward pass
* renaming from *hetero to *heteroAPI
Co-authored-by: Israt Nisa <nisisrat@amazon.com>
Co-authored-by: Da Zheng <zhengda1936@gmail.com>
* squeeze node labels in FraudDataset
* fix RLModule
* update results in README.md
* fix KeyError in full graph training
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>
* add word_ids and simplify
* simplify
* add word_ids to be removed later
* remove word_ids
* seems to work
* tweak
* transpose word_z
* add word_ids example
* check api compatibility
* improve compatibility
* update doc
* tweak verbose
* restore word_z layout; tweak
* tweak
* tweak doc
* word_cT
* use log_weight and some other tweaks
* rewrite README
* update equations
* rewrite for clarity and pass tests
* tweak
* bugfix import
* fix unit test
* fix mult to be the same as old versions
* tweak
* could be a bugfix
* 0/0=nan
* add doc_subgraph utility function
* minor cache optimization
* minor cache tweak
* add environmental variable to trade cache speed for memory
* update README
* tweak
* add sparse update pass unit test
* simplify sparse update
* improve low-memory efficiency
* tweak
* add sample expectation scores to allow resampling
* simplify
* update comment
* avoid edge cases
* bugfix pred scores
* simplify
* add save function
Co-authored-by: Yifei Ma <yifeim@amazon.com>
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
Co-authored-by: Jinjing Zhou <VoVAllen@users.noreply.github.com>