* Remove all torchtext legacy-related APIs
* Remove unused BagOfWordsPretrained class, and fix some typos
Co-authored-by: Mufei Li <mufeili1996@gmail.com>
* Explicitly unpin tensoradapter allocated arrays
* Undo unrelated change
* Add unit test
* update unit test
* add pinned_by_dgl flag to NDArray::Container
* use dgl.ndarray for holding the pinning status
* update multi-gpu uva inference
* reinterpret cast NDArray::Container* to DLTensor* in MoveAsDLTensor
* update unpin column and examples
* add unit test for unpin column
Co-authored-by: Dominique LaSalle <dlasalle@nvidia.com>
Co-authored-by: nv-dlasalle <63612878+nv-dlasalle@users.noreply.github.com>
* Fix a cub compile error for CUDA 11.5
* Fix comparison of integer expressions of different signedness in coo_sort.cu file
* Fix comparison of integer expressions of different signedness in cuda_compact_graph.cu file
* Remove never referenced variable in spmm.cu
* Fix comparison of integer expressions of different signedness in rowwise_pick.h file
* Fix comparison of integer expressions of different signedness in choice.cc file
* Remove never referenced variable col_data in spat_op_impl_coo.cc
* Remove never referenced variable allowed in global_uniform.cc
* Fix comparison of integer expressions of different signedness in graph.cc
* Fix comparison of integer expressions of different signedness in graph_apis.cc
* Fix the un-used ctx variable in ndarray_partition.cc file for cpu only build
* Fix comparison of integer expressions of different signedness in libra_partition.cc
* Fix comparison of integer expressions of different signedness in graph_op.cc
Co-authored-by: Triston Cao <tristonc@nvidia.com>
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
* fix unstable sort
* add torch version check
* reformat
* split too long comments
* Update dataloader.py
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
* * Added functions from dgl.transforms.functional that were missing from the list for documentation in dgl.rst
* * Sorted transform ops list in dgl.rst in alphabetical order
Co-authored-by: Xin Yao <xiny@nvidia.com>
* Fix fail to create_shared_mem_array in ddp spawn train #4110
Fix fail to create_shared_mem_array in ddp spawn train #4110
* [Bugfix] Fix fail to create_shared_mem_array in ddp spawn train #4110
[Bugfix] Fix fail to create_shared_mem_array in ddp spawn train #4110
Replace random.seed() to random_ = random.Random()
* Update pytorch.py
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
* add argument reorder=False for citation_graph
* add description of the argument reorder
* add reordered/un_reordered save_path
* add version number postfix
Co-authored-by: Mufei Li <mufeili1996@gmail.com>
* Wrap all CUDA runtime API/CUB calls with macro
* remove the usage of explicit cudaMalloc in favor of AllocWorkspace
* fix typo
Co-authored-by: Israt Nisa <neesha295@gmail.com>
* [DistTest] add basic pipeline for dist test across machines
* move launch remote cmd to separate file
* add test for rpc
* fix function naming rule
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
* * Added specialization of cublasGemm function for `__half` type, to try to address https://github.com/dmlc/dgl/issues/3988
* * Added USE_FP16 guard
* * Added test cases to test_segment_mm, to test newly-added FP16 specialization of cublasGemm
* * Replaced for loop in test_segment_mm with pytest.mark.parametrize, as recommended
Co-authored-by: Xin Yao <xiny@nvidia.com>
* * Added support for common operations on FP16 (`half` or `__half`) for older GPU architectures
* Fixed an issue with previous check for FP16 support
* * Removing FP16 type checks, since they should no longer be needed
* * Fixed AtomicAdd to be atomic for `float` and `double` for old GPU architectures. Unfortunately, it seems that atomicCAS for unsigned short seems to be unavailable until architecture 70, so half will have to stay non-atomic on old GPUs.
* * Fixed non-atomic version of `AtomicAdd<half>` for older GPUs to return old value instead value of new