dmlc--dgl
e452179c88
* Update from master (#4584)
* [Example][Refactor] Refactor graphsage multigpu and full-graph example (#4430)
* Add refactors for multi-gpu and full-graph example
* Fix format
* Update
* Update
* Update
* [Cleanup] Remove async_transferer (#4505)
* Remove async_transferer
* remove test
* Remove AsyncTransferer
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: Xin Yao <yaox12@outlook.com>
* [Cleanup] Remove duplicate entries of CUB submodule (issue# 4395) (#4499)
* remove third_part/cub
* remove from third_party
Co-authored-by: Israt Nisa <nisisrat@amazon.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
* [Bug] Enable turn on/off libxsmm at runtime (#4455)
* enable turn on/off libxsmm at runtime by adding a global config and related API
Co-authored-by: Ubuntu <ubuntu@ip-172-31-19-194.ap-northeast-1.compute.internal>
* [Feature] Unify the cuda stream used in core library (#4480)
* Use an internal cuda stream for CopyDataFromTo
* small fix white space
* Fix to compile
* Make stream optional in copydata for compile
* fix lint issue
* Update cub functions to use internal stream
* Lint check
* Update CopyTo/CopyFrom/CopyFromTo to use internal stream
* Address comments
* Fix backward CUDA stream
* Avoid overloading CopyFromTo()
* Minor comment update
* Overload copydatafromto in cuda device api
Co-authored-by: xiny <xiny@nvidia.com>
* [Feature] Added exclude_self and output_batch to knn graph construction (Issues #4323 #4316) (#4389)
* * Added "exclude_self" and "output_batch" options to knn_graph and segmented_knn_graph
* Updated out-of-date comments on remove_edges and remove_self_loop, since they now preserve batch information
* * Changed defaults on new knn_graph and segmented_knn_graph function parameters, for compatibility; pytorch/test_geometry.py was failing
* * Added test to ensure dgl.remove_self_loop function correctly updates batch information
* * Added new knn_graph and segmented_knn_graph parameters to dgl.nn.KNNGraph and dgl.nn.SegmentedKNNGraph
* * Formatting
* * Oops, I missed the one in segmented_knn_graph when I fixed the similar thing in knn_graph
* * Fixed edge case handling when invalid k specified, since it still needs to be handled consistently for tests to pass
* Fixed context of batch info, since it must match the context of the input position data for remove_self_loop to succeed
* * Fixed batch info resulting from knn_graph when output_batch is true, for case of 3D input tensor, representing multiple segments
* * Added testing of new exclude_self and output_batch parameters on knn_graph and segmented_knn_graph, and their wrappers, KNNGraph and SegmentedKNNGraph, into the test_knn_cuda test
* * Added doc comments for new parameters
* * Added correct handling for uncommon case of k or more coincident points when excluding self edges in knn_graph and segmented_knn_graph
* Added test cases for more than k coincident points
* * Updated doc comments for output_batch parameters for clarity
* * Linter formatting fixes
* * Extracted out common function for test_knn_cpu and test_knn_cuda, to add the new test cases to test_knn_cpu
* * Rewording in doc comments
* * Removed output_batch parameter from knn_graph and segmented_knn_graph, in favour of always setting the batch information, except in knn_graph if x is a 2D tensor
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
* [CI] only known devs are authorized to trigger CI (#4518)
* [CI] only known devs are authorized to trigger CI
* fix if author is null
* add comments
* [Readability] Auto fix setup.py and update-version.py (#4446)
* Auto fix update-version
* Auto fix setup.py
* Auto fix update-version
* Auto fix setup.py
* [Doc] Change random.py to random_partition.py in guide on distributed partition pipeline (#4438)
* Update distributed-preprocessing.rst
* Update
Co-authored-by: Ubuntu <ubuntu@ip-172-31-9-26.ap-northeast-1.compute.internal>
* fix unpinning when tensoradaptor is not available (#4450)
* [Doc] fix print issue in tutorial (#4459)
* [Example][Refactor] Refactor RGCN example (#4327)
* Refactor full graph entity classification
* Refactor rgcn with sampling
* README update
* Update
* Results update
* Respect default setting of self_loop=false in entity.py
* Update
* Update README
* Update for multi-gpu
* Update
* [doc] fix invalid link in user guide (#4468)
* [Example] directional_GSN for ogbg-molpcba (#4405)
* version-1
* version-2
* version-3
* update examples/README
* Update .gitignore
* update performance in README, delete scripts
* 1st approving review
* 2nd approving review
Co-authored-by: Mufei Li <mufeili1996@gmail.com>
* Clarify the message name, which is 'm'. (#4462)
Co-authored-by: Ubuntu <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
Co-authored-by: Rhett Ying <85214957+Rhett-Ying@users.noreply.github.com>
* [Refactor] Auto fix view.py. (#4461)
Co-authored-by: Ubuntu <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
* [Example] SEAL for OGBL (#4291)
* [Example] SEAL for OGBL
* update index
* update
* fix readme typo
* add seal sampler
* modify set ops
* prefetch
* efficiency test
* update
* optimize
* fix ScatterAdd dtype issue
* update sampler style
* update
Co-authored-by: Quan Gan <coin2028@hotmail.com>
* [CI] use https instead of http (#4488)
* [BugFix] fix crash due to incorrect dtype in dgl.to_block() (#4487)
* [BugFix] fix crash due to incorrect dtype in dgl.to_block()
* fix test failure in TF
* [Feature] Make TensorAdapter Stream Aware (#4472)
* Allocate tensors in DGL's current stream
* make tensoradaptor stream-aware
* replace TAemtpy with cpu allocator
* fix typo
* try fix cpu allocation
* clean header
* redirect AllocDataSpace as well
* resolve comments
* [Build][Doc] Specify the sphinx version (#4465)
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
* reformat
* reformat
* Auto fix update-version
* Auto fix setup.py
* reformat
* reformat
Co-authored-by: Ubuntu <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
Co-authored-by: Rhett Ying <85214957+Rhett-Ying@users.noreply.github.com>
Co-authored-by: Mufei Li <mufeili1996@gmail.com>
Co-authored-by: Ubuntu <ubuntu@ip-172-31-9-26.ap-northeast-1.compute.internal>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: Chang Liu <chang.liu@utexas.edu>
Co-authored-by: Zhiteng Li <55398076+ZHITENGLI@users.noreply.github.com>
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
Co-authored-by: rudongyu <ru_dongyu@outlook.com>
Co-authored-by: Quan Gan <coin2028@hotmail.com>
* Move mock version of dgl_sparse library to DGL main repo (#4524)
* init
* Add api doc for sparse library
* support op btwn matrices with differnt sparsity
* Fixed docstring
* addresses comments
* lint check
* change keyword format to fmt
Co-authored-by: Israt Nisa <nisisrat@amazon.com>
* [DistPart] expose timeout config for process group (#4532)
* [DistPart] expose timeout config for process group
* refine code
* Update tools/distpartitioning/data_proc_pipeline.py
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
* [Feature] Import PyTorch's CUDA stream management (#4503)
* add set_stream
* add .record_stream for NDArray and HeteroGraph
* refactor dgl stream Python APIs
* test record_stream
* add unit test for record stream
* use pytorch's stream
* fix lint
* fix cpu build
* address comments
* address comments
* add record stream tests for dgl.graph
* record frames and update dataloder
* add docstring
* update frame
* add backend check for record_stream
* remove CUDAThreadEntry::stream
* record stream for newly created formats
* fix bug
* fix cpp test
* fix None c_void_p to c_handle
* [examples]educe memory consumption (#4558)
* [examples]educe memory consumption
* reffine help message
* refine
* [Feature][REVIEW] Enable DGL cugaph nightly CI (#4525)
* Added cugraph nightly scripts
* Removed nvcr.io//nvidia/pytorch:22.04-py3 reference
Co-authored-by: Rhett Ying <85214957+Rhett-Ying@users.noreply.github.com>
* Revert "[Feature][REVIEW] Enable DGL cugaph nightly CI (#4525)" (#4563)
This reverts commit ec171c648a.
* [Misc] Add flake8 lint workflow. (#4566)
* Add pyproject.toml for autopep8.
* Add pyproject.toml for autopep8.
* Add flake8 annotation in workflow.
* remove
* add
* clean up
Co-authored-by: Steve <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
* [Misc] Try use official pylint workflow. (#4568)
* polish update_version
* update pylint workflow.
* add
* revert.
Co-authored-by: Steve <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
* [CI] refine stage logic (#4565)
* [CI] refine stage logic
* refine
* refine
* remove (#4570)
Co-authored-by: Steve <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
* Add Pylint workflow for flake8. (#4571)
* remove
* Add pylint.
Co-authored-by: Steve <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
* [Misc] Update the python version in Pylint workflow for flake8. (#4572)
* remove
* Add pylint.
* Change the python version for pylint.
Co-authored-by: Steve <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
* Update pylint. (#4574)
Co-authored-by: Steve <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
* [Misc] Use another workflow. (#4575)
* Update pylint.
* Use another workflow.
Co-authored-by: Steve <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
* Update pylint. (#4576)
Co-authored-by: Steve <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
* Update pylint.yml
* Update pylint.yml
* Delete pylint.yml
* [Misc]Add pyproject.toml for autopep8 & black. (#4543)
* Add pyproject.toml for autopep8.
* Add pyproject.toml for autopep8.
Co-authored-by: Steve <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
* [Feature] Bump DLPack to v0.7 and decouple DLPack from the core library (#4454)
* rename `DLContext` to `DGLContext`
* rename `kDLGPU` to `kDLCUDA`
* replace DLTensor with DGLArray
* fix linting
* Unify DGLType and DLDataType to DGLDataType
* Fix FFI
* rename DLDeviceType to DGLDeviceType
* decouple dlpack from the core library
* fix bug
* fix lint
* fix merge
* fix build
* address comments
* rename dl_converter to dlpack_convert
* remove redundant comments
Co-authored-by: Chang Liu <chang.liu@utexas.edu>
Co-authored-by: nv-dlasalle <63612878+nv-dlasalle@users.noreply.github.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: Xin Yao <yaox12@outlook.com>
Co-authored-by: Israt Nisa <neesha295@gmail.com>
Co-authored-by: Israt Nisa <nisisrat@amazon.com>
Co-authored-by: peizhou001 <110809584+peizhou001@users.noreply.github.com>
Co-authored-by: Ubuntu <ubuntu@ip-172-31-19-194.ap-northeast-1.compute.internal>
Co-authored-by: ndickson-nvidia <99772994+ndickson-nvidia@users.noreply.github.com>
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
Co-authored-by: Rhett Ying <85214957+Rhett-Ying@users.noreply.github.com>
Co-authored-by: Hongzhi (Steve), Chen <chenhongzhi.nkcs@gmail.com>
Co-authored-by: Ubuntu <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
Co-authored-by: Ubuntu <ubuntu@ip-172-31-9-26.ap-northeast-1.compute.internal>
Co-authored-by: Zhiteng Li <55398076+ZHITENGLI@users.noreply.github.com>
Co-authored-by: rudongyu <ru_dongyu@outlook.com>
Co-authored-by: Quan Gan <coin2028@hotmail.com>
Co-authored-by: Vibhu Jawa <vibhujawa@gmail.com>
* [Deprecation] Dataset Attributes (#4546)
* Update
* CI
* CI
* Update
Co-authored-by: Ubuntu <ubuntu@ip-172-31-9-26.ap-northeast-1.compute.internal>
* [Example] Bug Fix (#4665)
* Update
* CI
* CI
* Update
* Update
Co-authored-by: Ubuntu <ubuntu@ip-172-31-9-26.ap-northeast-1.compute.internal>
* Update
Co-authored-by: Chang Liu <chang.liu@utexas.edu>
Co-authored-by: nv-dlasalle <63612878+nv-dlasalle@users.noreply.github.com>
Co-authored-by: Xin Yao <xiny@nvidia.com>
Co-authored-by: Xin Yao <yaox12@outlook.com>
Co-authored-by: Israt Nisa <neesha295@gmail.com>
Co-authored-by: Israt Nisa <nisisrat@amazon.com>
Co-authored-by: peizhou001 <110809584+peizhou001@users.noreply.github.com>
Co-authored-by: Ubuntu <ubuntu@ip-172-31-19-194.ap-northeast-1.compute.internal>
Co-authored-by: ndickson-nvidia <99772994+ndickson-nvidia@users.noreply.github.com>
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
Co-authored-by: Rhett Ying <85214957+Rhett-Ying@users.noreply.github.com>
Co-authored-by: Hongzhi (Steve), Chen <chenhongzhi.nkcs@gmail.com>
Co-authored-by: Ubuntu <ubuntu@ip-172-31-34-29.ap-northeast-1.compute.internal>
Co-authored-by: Ubuntu <ubuntu@ip-172-31-9-26.ap-northeast-1.compute.internal>
Co-authored-by: Zhiteng Li <55398076+ZHITENGLI@users.noreply.github.com>
Co-authored-by: rudongyu <ru_dongyu@outlook.com>
Co-authored-by: Quan Gan <coin2028@hotmail.com>
Co-authored-by: Vibhu Jawa <vibhujawa@gmail.com>
814 行
28 KiB
Python
814 行
28 KiB
Python
"""Cora, citeseer, pubmed dataset.
|
|
|
|
(lingfan): following dataset loading and preprocessing code from tkipf/gcn
|
|
https://github.com/tkipf/gcn/blob/master/gcn/utils.py
|
|
"""
|
|
from __future__ import absolute_import
|
|
|
|
import numpy as np
|
|
import pickle as pkl
|
|
import networkx as nx
|
|
import scipy.sparse as sp
|
|
import os, sys
|
|
|
|
from .utils import save_graphs, load_graphs, save_info, load_info, makedirs, _get_dgl_url
|
|
from .utils import generate_mask_tensor
|
|
from .utils import deprecate_property, deprecate_function
|
|
from .dgl_dataset import DGLBuiltinDataset
|
|
from .. import convert
|
|
from .. import batch
|
|
from .. import backend as F
|
|
from ..convert import graph as dgl_graph
|
|
from ..convert import from_networkx, to_networkx
|
|
from ..transforms import reorder_graph
|
|
|
|
backend = os.environ.get('DGLBACKEND', 'pytorch')
|
|
|
|
def _pickle_load(pkl_file):
|
|
if sys.version_info > (3, 0):
|
|
return pkl.load(pkl_file, encoding='latin1')
|
|
else:
|
|
return pkl.load(pkl_file)
|
|
|
|
class CitationGraphDataset(DGLBuiltinDataset):
|
|
r"""The citation graph dataset, including cora, citeseer and pubmeb.
|
|
Nodes mean authors and edges mean citation relationships.
|
|
|
|
Parameters
|
|
-----------
|
|
name: str
|
|
name can be 'cora', 'citeseer' or 'pubmed'.
|
|
raw_dir : str
|
|
Raw file directory to download/contains the input data directory.
|
|
Default: ~/.dgl/
|
|
force_reload : bool
|
|
Whether to reload the dataset. Default: False
|
|
verbose : bool
|
|
Whether to print out progress information. Default: True.
|
|
reverse_edge : bool
|
|
Whether to add reverse edges in graph. Default: True.
|
|
transform : callable, optional
|
|
A transform that takes in a :class:`~dgl.DGLGraph` object and returns
|
|
a transformed version. The :class:`~dgl.DGLGraph` object will be
|
|
transformed before every access.
|
|
reorder : bool
|
|
Whether to reorder the graph using :func:`~dgl.reorder_graph`. Default: False.
|
|
"""
|
|
_urls = {
|
|
'cora_v2' : 'dataset/cora_v2.zip',
|
|
'citeseer' : 'dataset/citeseer.zip',
|
|
'pubmed' : 'dataset/pubmed.zip',
|
|
}
|
|
|
|
def __init__(self, name, raw_dir=None, force_reload=False,
|
|
verbose=True, reverse_edge=True, transform=None,
|
|
reorder=False):
|
|
assert name.lower() in ['cora', 'citeseer', 'pubmed']
|
|
|
|
# Previously we use the pre-processing in pygcn (https://github.com/tkipf/pygcn)
|
|
# for Cora, which is slightly different from the one used in the GCN paper
|
|
if name.lower() == 'cora':
|
|
name = 'cora_v2'
|
|
|
|
url = _get_dgl_url(self._urls[name])
|
|
self._reverse_edge = reverse_edge
|
|
self._reorder = reorder
|
|
|
|
super(CitationGraphDataset, self).__init__(name,
|
|
url=url,
|
|
raw_dir=raw_dir,
|
|
force_reload=force_reload,
|
|
verbose=verbose,
|
|
transform=transform)
|
|
|
|
def process(self):
|
|
"""Loads input data from data directory and reorder graph for better locality
|
|
|
|
ind.name.x => the feature vectors of the training instances as scipy.sparse.csr.csr_matrix object;
|
|
ind.name.tx => the feature vectors of the test instances as scipy.sparse.csr.csr_matrix object;
|
|
ind.name.allx => the feature vectors of both labeled and unlabeled training instances
|
|
(a superset of ind.name.x) as scipy.sparse.csr.csr_matrix object;
|
|
ind.name.y => the one-hot labels of the labeled training instances as numpy.ndarray object;
|
|
ind.name.ty => the one-hot labels of the test instances as numpy.ndarray object;
|
|
ind.name.ally => the labels for instances in ind.name.allx as numpy.ndarray object;
|
|
ind.name.graph => a dict in the format {index: [index_of_neighbor_nodes]} as collections.defaultdict
|
|
object;
|
|
ind.name.test.index => the indices of test instances in graph, for the inductive setting as list object.
|
|
"""
|
|
root = self.raw_path
|
|
objnames = ['x', 'y', 'tx', 'ty', 'allx', 'ally', 'graph']
|
|
objects = []
|
|
for i in range(len(objnames)):
|
|
with open("{}/ind.{}.{}".format(root, self.name, objnames[i]), 'rb') as f:
|
|
objects.append(_pickle_load(f))
|
|
|
|
x, y, tx, ty, allx, ally, graph = tuple(objects)
|
|
test_idx_reorder = _parse_index_file("{}/ind.{}.test.index".format(root, self.name))
|
|
test_idx_range = np.sort(test_idx_reorder)
|
|
|
|
if self.name == 'citeseer':
|
|
# Fix citeseer dataset (there are some isolated nodes in the graph)
|
|
# Find isolated nodes, add them as zero-vecs into the right position
|
|
test_idx_range_full = range(min(test_idx_reorder), max(test_idx_reorder)+1)
|
|
tx_extended = sp.lil_matrix((len(test_idx_range_full), x.shape[1]))
|
|
tx_extended[test_idx_range-min(test_idx_range), :] = tx
|
|
tx = tx_extended
|
|
ty_extended = np.zeros((len(test_idx_range_full), y.shape[1]))
|
|
ty_extended[test_idx_range-min(test_idx_range), :] = ty
|
|
ty = ty_extended
|
|
|
|
features = sp.vstack((allx, tx)).tolil()
|
|
features[test_idx_reorder, :] = features[test_idx_range, :]
|
|
|
|
if self.reverse_edge:
|
|
graph = nx.DiGraph(nx.from_dict_of_lists(graph))
|
|
g = from_networkx(graph)
|
|
else:
|
|
graph = nx.Graph(nx.from_dict_of_lists(graph))
|
|
edges = list(graph.edges())
|
|
u, v = map(list, zip(*edges))
|
|
g = dgl_graph((u, v))
|
|
|
|
onehot_labels = np.vstack((ally, ty))
|
|
onehot_labels[test_idx_reorder, :] = onehot_labels[test_idx_range, :]
|
|
labels = np.argmax(onehot_labels, 1)
|
|
|
|
idx_test = test_idx_range.tolist()
|
|
idx_train = range(len(y))
|
|
idx_val = range(len(y), len(y)+500)
|
|
|
|
train_mask = generate_mask_tensor(_sample_mask(idx_train, labels.shape[0]))
|
|
val_mask = generate_mask_tensor(_sample_mask(idx_val, labels.shape[0]))
|
|
test_mask = generate_mask_tensor(_sample_mask(idx_test, labels.shape[0]))
|
|
|
|
g.ndata['train_mask'] = train_mask
|
|
g.ndata['val_mask'] = val_mask
|
|
g.ndata['test_mask'] = test_mask
|
|
g.ndata['label'] = F.tensor(labels)
|
|
g.ndata['feat'] = F.tensor(_preprocess_features(features), dtype=F.data_type_dict['float32'])
|
|
self._num_classes = onehot_labels.shape[1]
|
|
self._labels = labels
|
|
if self._reorder:
|
|
self._g = reorder_graph(
|
|
g, node_permute_algo='rcmk', edge_permute_algo='dst', store_ids=False)
|
|
else:
|
|
self._g = g
|
|
|
|
if self.verbose:
|
|
print('Finished data loading and preprocessing.')
|
|
print(' NumNodes: {}'.format(self._g.number_of_nodes()))
|
|
print(' NumEdges: {}'.format(self._g.number_of_edges()))
|
|
print(' NumFeats: {}'.format(self._g.ndata['feat'].shape[1]))
|
|
print(' NumClasses: {}'.format(self.num_classes))
|
|
print(' NumTrainingSamples: {}'.format(
|
|
F.nonzero_1d(self._g.ndata['train_mask']).shape[0]))
|
|
print(' NumValidationSamples: {}'.format(
|
|
F.nonzero_1d(self._g.ndata['val_mask']).shape[0]))
|
|
print(' NumTestSamples: {}'.format(
|
|
F.nonzero_1d(self._g.ndata['test_mask']).shape[0]))
|
|
|
|
def has_cache(self):
|
|
graph_path = os.path.join(self.save_path,
|
|
self.save_name + '.bin')
|
|
info_path = os.path.join(self.save_path,
|
|
self.save_name + '.pkl')
|
|
if os.path.exists(graph_path) and \
|
|
os.path.exists(info_path):
|
|
return True
|
|
|
|
return False
|
|
|
|
def save(self):
|
|
"""save the graph list and the labels"""
|
|
graph_path = os.path.join(self.save_path,
|
|
self.save_name + '.bin')
|
|
info_path = os.path.join(self.save_path,
|
|
self.save_name + '.pkl')
|
|
save_graphs(str(graph_path), self._g)
|
|
save_info(str(info_path), {'num_classes': self.num_classes})
|
|
|
|
def load(self):
|
|
graph_path = os.path.join(self.save_path,
|
|
self.save_name + '.bin')
|
|
info_path = os.path.join(self.save_path,
|
|
self.save_name + '.pkl')
|
|
graphs, _ = load_graphs(str(graph_path))
|
|
|
|
info = load_info(str(info_path))
|
|
graph = graphs[0]
|
|
self._g = graph
|
|
# for compatability
|
|
graph = graph.clone()
|
|
graph.ndata.pop('train_mask')
|
|
graph.ndata.pop('val_mask')
|
|
graph.ndata.pop('test_mask')
|
|
graph.ndata.pop('feat')
|
|
graph.ndata.pop('label')
|
|
graph = to_networkx(graph)
|
|
|
|
self._num_classes = info['num_classes']
|
|
self._g.ndata['train_mask'] = generate_mask_tensor(F.asnumpy(self._g.ndata['train_mask']))
|
|
self._g.ndata['val_mask'] = generate_mask_tensor(F.asnumpy(self._g.ndata['val_mask']))
|
|
self._g.ndata['test_mask'] = generate_mask_tensor(F.asnumpy(self._g.ndata['test_mask']))
|
|
# hack for mxnet compatability
|
|
|
|
if self.verbose:
|
|
print(' NumNodes: {}'.format(self._g.number_of_nodes()))
|
|
print(' NumEdges: {}'.format(self._g.number_of_edges()))
|
|
print(' NumFeats: {}'.format(self._g.ndata['feat'].shape[1]))
|
|
print(' NumClasses: {}'.format(self.num_classes))
|
|
print(' NumTrainingSamples: {}'.format(
|
|
F.nonzero_1d(self._g.ndata['train_mask']).shape[0]))
|
|
print(' NumValidationSamples: {}'.format(
|
|
F.nonzero_1d(self._g.ndata['val_mask']).shape[0]))
|
|
print(' NumTestSamples: {}'.format(
|
|
F.nonzero_1d(self._g.ndata['test_mask']).shape[0]))
|
|
|
|
def __getitem__(self, idx):
|
|
assert idx == 0, "This dataset has only one graph"
|
|
if self._transform is None:
|
|
return self._g
|
|
else:
|
|
return self._transform(self._g)
|
|
|
|
def __len__(self):
|
|
return 1
|
|
|
|
@property
|
|
def save_name(self):
|
|
return self.name + '_dgl_graph'
|
|
|
|
@property
|
|
def num_labels(self):
|
|
deprecate_property('dataset.num_labels', 'dataset.num_classes')
|
|
return self.num_classes
|
|
|
|
@property
|
|
def num_classes(self):
|
|
return self._num_classes
|
|
|
|
""" Citation graph is used in many examples
|
|
We preserve these properties for compatability.
|
|
"""
|
|
|
|
@property
|
|
def reverse_edge(self):
|
|
return self._reverse_edge
|
|
|
|
|
|
def _preprocess_features(features):
|
|
"""Row-normalize feature matrix and convert to tuple representation"""
|
|
rowsum = np.asarray(features.sum(1))
|
|
r_inv = np.power(rowsum, -1).flatten()
|
|
r_inv[np.isinf(r_inv)] = 0.
|
|
r_mat_inv = sp.diags(r_inv)
|
|
features = r_mat_inv.dot(features)
|
|
return np.asarray(features.todense())
|
|
|
|
def _parse_index_file(filename):
|
|
"""Parse index file."""
|
|
index = []
|
|
for line in open(filename):
|
|
index.append(int(line.strip()))
|
|
return index
|
|
|
|
def _sample_mask(idx, l):
|
|
"""Create mask."""
|
|
mask = np.zeros(l)
|
|
mask[idx] = 1
|
|
return mask
|
|
|
|
class CoraGraphDataset(CitationGraphDataset):
|
|
r""" Cora citation network dataset.
|
|
|
|
Nodes mean paper and edges mean citation
|
|
relationships. Each node has a predefined
|
|
feature with 1433 dimensions. The dataset is
|
|
designed for the node classification task.
|
|
The task is to predict the category of
|
|
certain paper.
|
|
|
|
Statistics:
|
|
|
|
- Nodes: 2708
|
|
- Edges: 10556
|
|
- Number of Classes: 7
|
|
- Label split:
|
|
|
|
- Train: 140
|
|
- Valid: 500
|
|
- Test: 1000
|
|
|
|
Parameters
|
|
----------
|
|
raw_dir : str
|
|
Raw file directory to download/contains the input data directory.
|
|
Default: ~/.dgl/
|
|
force_reload : bool
|
|
Whether to reload the dataset. Default: False
|
|
verbose : bool
|
|
Whether to print out progress information. Default: True.
|
|
reverse_edge : bool
|
|
Whether to add reverse edges in graph. Default: True.
|
|
transform : callable, optional
|
|
A transform that takes in a :class:`~dgl.DGLGraph` object and returns
|
|
a transformed version. The :class:`~dgl.DGLGraph` object will be
|
|
transformed before every access.
|
|
reorder : bool
|
|
Whether to reorder the graph using :func:`~dgl.reorder_graph`. Default: False.
|
|
|
|
Attributes
|
|
----------
|
|
num_classes: int
|
|
Number of label classes
|
|
|
|
Notes
|
|
-----
|
|
The node feature is row-normalized.
|
|
|
|
Examples
|
|
--------
|
|
>>> dataset = CoraGraphDataset()
|
|
>>> g = dataset[0]
|
|
>>> num_class = dataset.num_classes
|
|
>>>
|
|
>>> # get node feature
|
|
>>> feat = g.ndata['feat']
|
|
>>>
|
|
>>> # get data split
|
|
>>> train_mask = g.ndata['train_mask']
|
|
>>> val_mask = g.ndata['val_mask']
|
|
>>> test_mask = g.ndata['test_mask']
|
|
>>>
|
|
>>> # get labels
|
|
>>> label = g.ndata['label']
|
|
|
|
"""
|
|
def __init__(self, raw_dir=None, force_reload=False, verbose=True,
|
|
reverse_edge=True, transform=None, reorder=False):
|
|
name = 'cora'
|
|
|
|
super(CoraGraphDataset, self).__init__(name, raw_dir, force_reload,
|
|
verbose, reverse_edge, transform, reorder)
|
|
|
|
def __getitem__(self, idx):
|
|
r"""Gets the graph object
|
|
|
|
Parameters
|
|
-----------
|
|
idx: int
|
|
Item index, CoraGraphDataset has only one graph object
|
|
|
|
Return
|
|
------
|
|
:class:`dgl.DGLGraph`
|
|
|
|
graph structure, node features and labels.
|
|
|
|
- ``ndata['train_mask']``: mask for training node set
|
|
- ``ndata['val_mask']``: mask for validation node set
|
|
- ``ndata['test_mask']``: mask for test node set
|
|
- ``ndata['feat']``: node feature
|
|
- ``ndata['label']``: ground truth labels
|
|
"""
|
|
return super(CoraGraphDataset, self).__getitem__(idx)
|
|
|
|
def __len__(self):
|
|
r"""The number of graphs in the dataset."""
|
|
return super(CoraGraphDataset, self).__len__()
|
|
|
|
class CiteseerGraphDataset(CitationGraphDataset):
|
|
r""" Citeseer citation network dataset.
|
|
|
|
Nodes mean scientific publications and edges
|
|
mean citation relationships. Each node has a
|
|
predefined feature with 3703 dimensions. The
|
|
dataset is designed for the node classification
|
|
task. The task is to predict the category of
|
|
certain publication.
|
|
|
|
Statistics:
|
|
|
|
- Nodes: 3327
|
|
- Edges: 9228
|
|
- Number of Classes: 6
|
|
- Label Split:
|
|
|
|
- Train: 120
|
|
- Valid: 500
|
|
- Test: 1000
|
|
|
|
Parameters
|
|
-----------
|
|
raw_dir : str
|
|
Raw file directory to download/contains the input data directory.
|
|
Default: ~/.dgl/
|
|
force_reload : bool
|
|
Whether to reload the dataset. Default: False
|
|
verbose : bool
|
|
Whether to print out progress information. Default: True.
|
|
reverse_edge : bool
|
|
Whether to add reverse edges in graph. Default: True.
|
|
transform : callable, optional
|
|
A transform that takes in a :class:`~dgl.DGLGraph` object and returns
|
|
a transformed version. The :class:`~dgl.DGLGraph` object will be
|
|
transformed before every access.
|
|
reorder : bool
|
|
Whether to reorder the graph using :func:`~dgl.reorder_graph`. Default: False.
|
|
|
|
Attributes
|
|
----------
|
|
num_classes: int
|
|
Number of label classes
|
|
|
|
Notes
|
|
-----
|
|
The node feature is row-normalized.
|
|
|
|
In citeseer dataset, there are some isolated nodes in the graph.
|
|
These isolated nodes are added as zero-vecs into the right position.
|
|
|
|
Examples
|
|
--------
|
|
>>> dataset = CiteseerGraphDataset()
|
|
>>> g = dataset[0]
|
|
>>> num_class = dataset.num_classes
|
|
>>>
|
|
>>> # get node feature
|
|
>>> feat = g.ndata['feat']
|
|
>>>
|
|
>>> # get data split
|
|
>>> train_mask = g.ndata['train_mask']
|
|
>>> val_mask = g.ndata['val_mask']
|
|
>>> test_mask = g.ndata['test_mask']
|
|
>>>
|
|
>>> # get labels
|
|
>>> label = g.ndata['label']
|
|
|
|
"""
|
|
def __init__(self, raw_dir=None, force_reload=False,
|
|
verbose=True, reverse_edge=True, transform=None, reorder=False):
|
|
name = 'citeseer'
|
|
|
|
super(CiteseerGraphDataset, self).__init__(name, raw_dir, force_reload,
|
|
verbose, reverse_edge, transform, reorder)
|
|
|
|
def __getitem__(self, idx):
|
|
r"""Gets the graph object
|
|
|
|
Parameters
|
|
-----------
|
|
idx: int
|
|
Item index, CiteseerGraphDataset has only one graph object
|
|
|
|
Return
|
|
------
|
|
:class:`dgl.DGLGraph`
|
|
|
|
graph structure, node features and labels.
|
|
|
|
- ``ndata['train_mask']``: mask for training node set
|
|
- ``ndata['val_mask']``: mask for validation node set
|
|
- ``ndata['test_mask']``: mask for test node set
|
|
- ``ndata['feat']``: node feature
|
|
- ``ndata['label']``: ground truth labels
|
|
"""
|
|
return super(CiteseerGraphDataset, self).__getitem__(idx)
|
|
|
|
def __len__(self):
|
|
r"""The number of graphs in the dataset."""
|
|
return super(CiteseerGraphDataset, self).__len__()
|
|
|
|
class PubmedGraphDataset(CitationGraphDataset):
|
|
r""" Pubmed citation network dataset.
|
|
|
|
Nodes mean scientific publications and edges
|
|
mean citation relationships. Each node has a
|
|
predefined feature with 500 dimensions. The
|
|
dataset is designed for the node classification
|
|
task. The task is to predict the category of
|
|
certain publication.
|
|
|
|
Statistics:
|
|
|
|
- Nodes: 19717
|
|
- Edges: 88651
|
|
- Number of Classes: 3
|
|
- Label Split:
|
|
|
|
- Train: 60
|
|
- Valid: 500
|
|
- Test: 1000
|
|
|
|
Parameters
|
|
-----------
|
|
raw_dir : str
|
|
Raw file directory to download/contains the input data directory.
|
|
Default: ~/.dgl/
|
|
force_reload : bool
|
|
Whether to reload the dataset. Default: False
|
|
verbose : bool
|
|
Whether to print out progress information. Default: True.
|
|
reverse_edge : bool
|
|
Whether to add reverse edges in graph. Default: True.
|
|
transform : callable, optional
|
|
A transform that takes in a :class:`~dgl.DGLGraph` object and returns
|
|
a transformed version. The :class:`~dgl.DGLGraph` object will be
|
|
transformed before every access.
|
|
reorder : bool
|
|
Whether to reorder the graph using :func:`~dgl.reorder_graph`. Default: False.
|
|
|
|
Attributes
|
|
----------
|
|
num_classes: int
|
|
Number of label classes
|
|
|
|
Notes
|
|
-----
|
|
The node feature is row-normalized.
|
|
|
|
Examples
|
|
--------
|
|
>>> dataset = PubmedGraphDataset()
|
|
>>> g = dataset[0]
|
|
>>> num_class = dataset.num_of_class
|
|
>>>
|
|
>>> # get node feature
|
|
>>> feat = g.ndata['feat']
|
|
>>>
|
|
>>> # get data split
|
|
>>> train_mask = g.ndata['train_mask']
|
|
>>> val_mask = g.ndata['val_mask']
|
|
>>> test_mask = g.ndata['test_mask']
|
|
>>>
|
|
>>> # get labels
|
|
>>> label = g.ndata['label']
|
|
|
|
"""
|
|
def __init__(self, raw_dir=None, force_reload=False, verbose=True,
|
|
reverse_edge=True, transform=None, reorder=False):
|
|
name = 'pubmed'
|
|
|
|
super(PubmedGraphDataset, self).__init__(name, raw_dir, force_reload,
|
|
verbose, reverse_edge, transform, reorder)
|
|
|
|
def __getitem__(self, idx):
|
|
r"""Gets the graph object
|
|
|
|
Parameters
|
|
-----------
|
|
idx: int
|
|
Item index, PubmedGraphDataset has only one graph object
|
|
|
|
Return
|
|
------
|
|
:class:`dgl.DGLGraph`
|
|
|
|
graph structure, node features and labels.
|
|
|
|
- ``ndata['train_mask']``: mask for training node set
|
|
- ``ndata['val_mask']``: mask for validation node set
|
|
- ``ndata['test_mask']``: mask for test node set
|
|
- ``ndata['feat']``: node feature
|
|
- ``ndata['label']``: ground truth labels
|
|
"""
|
|
return super(PubmedGraphDataset, self).__getitem__(idx)
|
|
|
|
def __len__(self):
|
|
r"""The number of graphs in the dataset."""
|
|
return super(PubmedGraphDataset, self).__len__()
|
|
|
|
def load_cora(raw_dir=None, force_reload=False, verbose=True, reverse_edge=True, transform=None):
|
|
"""Get CoraGraphDataset
|
|
|
|
Parameters
|
|
-----------
|
|
raw_dir : str
|
|
Raw file directory to download/contains the input data directory.
|
|
Default: ~/.dgl/
|
|
force_reload : bool
|
|
Whether to reload the dataset. Default: False
|
|
verbose : bool
|
|
Whether to print out progress information. Default: True.
|
|
reverse_edge : bool
|
|
Whether to add reverse edges in graph. Default: True.
|
|
transform : callable, optional
|
|
A transform that takes in a :class:`~dgl.DGLGraph` object and returns
|
|
a transformed version. The :class:`~dgl.DGLGraph` object will be
|
|
transformed before every access.
|
|
|
|
Return
|
|
-------
|
|
CoraGraphDataset
|
|
"""
|
|
data = CoraGraphDataset(raw_dir, force_reload, verbose, reverse_edge, transform)
|
|
return data
|
|
|
|
def load_citeseer(raw_dir=None, force_reload=False, verbose=True,
|
|
reverse_edge=True, transform=None):
|
|
"""Get CiteseerGraphDataset
|
|
|
|
Parameters
|
|
-----------
|
|
raw_dir : str
|
|
Raw file directory to download/contains the input data directory.
|
|
Default: ~/.dgl/
|
|
force_reload : bool
|
|
Whether to reload the dataset. Default: False
|
|
verbose : bool
|
|
Whether to print out progress information. Default: True.
|
|
reverse_edge : bool
|
|
Whether to add reverse edges in graph. Default: True.
|
|
transform : callable, optional
|
|
A transform that takes in a :class:`~dgl.DGLGraph` object and returns
|
|
a transformed version. The :class:`~dgl.DGLGraph` object will be
|
|
transformed before every access.
|
|
|
|
Return
|
|
-------
|
|
CiteseerGraphDataset
|
|
"""
|
|
data = CiteseerGraphDataset(raw_dir, force_reload, verbose, reverse_edge, transform)
|
|
return data
|
|
|
|
def load_pubmed(raw_dir=None, force_reload=False, verbose=True,
|
|
reverse_edge=True, transform=None):
|
|
"""Get PubmedGraphDataset
|
|
|
|
Parameters
|
|
-----------
|
|
raw_dir : str
|
|
Raw file directory to download/contains the input data directory.
|
|
Default: ~/.dgl/
|
|
force_reload : bool
|
|
Whether to reload the dataset. Default: False
|
|
verbose : bool
|
|
Whether to print out progress information. Default: True.
|
|
reverse_edge : bool
|
|
Whether to add reverse edges in graph. Default: True.
|
|
transform : callable, optional
|
|
A transform that takes in a :class:`~dgl.DGLGraph` object and returns
|
|
a transformed version. The :class:`~dgl.DGLGraph` object will be
|
|
transformed before every access.
|
|
|
|
Return
|
|
-------
|
|
PubmedGraphDataset
|
|
"""
|
|
data = PubmedGraphDataset(raw_dir, force_reload, verbose, reverse_edge, transform)
|
|
return data
|
|
|
|
class CoraBinary(DGLBuiltinDataset):
|
|
"""A mini-dataset for binary classification task using Cora.
|
|
|
|
After loaded, it has following members:
|
|
|
|
graphs : list of :class:`~dgl.DGLGraph`
|
|
pmpds : list of :class:`scipy.sparse.coo_matrix`
|
|
labels : list of :class:`numpy.ndarray`
|
|
|
|
Parameters
|
|
-----------
|
|
raw_dir : str
|
|
Raw file directory to download/contains the input data directory.
|
|
Default: ~/.dgl/
|
|
force_reload : bool
|
|
Whether to reload the dataset. Default: False
|
|
verbose: bool
|
|
Whether to print out progress information. Default: True.
|
|
transform : callable, optional
|
|
A transform that takes in a :class:`~dgl.DGLGraph` object and returns
|
|
a transformed version. The :class:`~dgl.DGLGraph` object will be
|
|
transformed before every access.
|
|
"""
|
|
def __init__(self, raw_dir=None, force_reload=False, verbose=True, transform=None):
|
|
name = 'cora_binary'
|
|
url = _get_dgl_url('dataset/cora_binary.zip')
|
|
super(CoraBinary, self).__init__(name,
|
|
url=url,
|
|
raw_dir=raw_dir,
|
|
force_reload=force_reload,
|
|
verbose=verbose,
|
|
transform=transform)
|
|
|
|
def process(self):
|
|
root = self.raw_path
|
|
# load graphs
|
|
self.graphs = []
|
|
with open("{}/graphs.txt".format(root), 'r') as f:
|
|
elist = []
|
|
for line in f.readlines():
|
|
if line.startswith('graph'):
|
|
if len(elist) != 0:
|
|
self.graphs.append(dgl_graph(tuple(zip(*elist))))
|
|
elist = []
|
|
else:
|
|
u, v = line.strip().split(' ')
|
|
elist.append((int(u), int(v)))
|
|
if len(elist) != 0:
|
|
self.graphs.append(dgl_graph(tuple(zip(*elist))))
|
|
with open("{}/pmpds.pkl".format(root), 'rb') as f:
|
|
self.pmpds = _pickle_load(f)
|
|
self.labels = []
|
|
with open("{}/labels.txt".format(root), 'r') as f:
|
|
cur = []
|
|
for line in f.readlines():
|
|
if line.startswith('graph'):
|
|
if len(cur) != 0:
|
|
self.labels.append(np.asarray(cur))
|
|
cur = []
|
|
else:
|
|
cur.append(int(line.strip()))
|
|
if len(cur) != 0:
|
|
self.labels.append(np.asarray(cur))
|
|
# sanity check
|
|
assert len(self.graphs) == len(self.pmpds)
|
|
assert len(self.graphs) == len(self.labels)
|
|
|
|
def has_cache(self):
|
|
graph_path = os.path.join(self.save_path,
|
|
self.save_name + '.bin')
|
|
if os.path.exists(graph_path):
|
|
return True
|
|
|
|
return False
|
|
|
|
def save(self):
|
|
"""save the graph list and the labels"""
|
|
graph_path = os.path.join(self.save_path,
|
|
self.save_name + '.bin')
|
|
labels = {}
|
|
for i, label in enumerate(self.labels):
|
|
labels['{}'.format(i)] = F.tensor(label)
|
|
save_graphs(str(graph_path), self.graphs, labels)
|
|
if self.verbose:
|
|
print('Done saving data into cached files.')
|
|
|
|
def load(self):
|
|
graph_path = os.path.join(self.save_path,
|
|
self.save_name + '.bin')
|
|
self.graphs, labels = load_graphs(str(graph_path))
|
|
|
|
self.labels = []
|
|
for i in range(len(labels)):
|
|
self.labels.append(F.asnumpy(labels['{}'.format(i)]))
|
|
# load pmpds under self.raw_path
|
|
with open("{}/pmpds.pkl".format(self.raw_path), 'rb') as f:
|
|
self.pmpds = _pickle_load(f)
|
|
if self.verbose:
|
|
print('Done loading data into cached files.')
|
|
# sanity check
|
|
assert len(self.graphs) == len(self.pmpds)
|
|
assert len(self.graphs) == len(self.labels)
|
|
|
|
def __len__(self):
|
|
return len(self.graphs)
|
|
|
|
def __getitem__(self, i):
|
|
r"""Gets the idx-th sample.
|
|
|
|
Parameters
|
|
-----------
|
|
idx : int
|
|
The sample index.
|
|
|
|
Returns
|
|
-------
|
|
(dgl.DGLGraph, scipy.sparse.coo_matrix, int)
|
|
The graph, scipy sparse coo_matrix and its label.
|
|
"""
|
|
if self._transform is None:
|
|
g = self.graphs[i]
|
|
else:
|
|
g = self._transform(self.graphs[i])
|
|
return (g, self.pmpds[i], self.labels[i])
|
|
|
|
@property
|
|
def save_name(self):
|
|
return self.name + '_dgl_graph'
|
|
|
|
@staticmethod
|
|
def collate_fn(cur):
|
|
graphs, pmpds, labels = zip(*cur)
|
|
batched_graphs = batch.batch(graphs)
|
|
batched_pmpds = sp.block_diag(pmpds)
|
|
batched_labels = np.concatenate(labels, axis=0)
|
|
return batched_graphs, batched_pmpds, batched_labels
|
|
|
|
def _normalize(mx):
|
|
"""Row-normalize sparse matrix"""
|
|
rowsum = np.asarray(mx.sum(1))
|
|
r_inv = np.power(rowsum, -1).flatten()
|
|
r_inv[np.isinf(r_inv)] = 0.
|
|
r_mat_inv = sp.diags(r_inv)
|
|
mx = r_mat_inv.dot(mx)
|
|
return mx
|
|
|
|
def _encode_onehot(labels):
|
|
classes = list(sorted(set(labels)))
|
|
classes_dict = {c: np.identity(len(classes))[i, :] for i, c in
|
|
enumerate(classes)}
|
|
labels_onehot = np.asarray(list(map(classes_dict.get, labels)),
|
|
dtype=np.int32)
|
|
return labels_onehot
|