dmlc--dgl
b0a9d16f25
* [Feature] Add full graph training with dgl built-in dataset. * [Feature] Add full graph training with dgl built-in dataset. * [Feature] Add full graph training with dgl built-in dataset. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Bug] fix model to cuda. * [Feature] Add test loss and accuracy * [Feature] Add test loss and accuracy * [Feature] Add test loss and accuracy * [Feature] Add test loss and accuracy * [Feature] Add test loss and accuracy * [Feature] Add test loss and accuracy * [Fix] Add random * [Bug] Fix batch norm error * [Doc] Test with CN in Sphinx * [Doc] Test with CN in Sphinx * [Doc] Remove the test CN docs. * [Feature] Add input embedding layer * [Feature] Add input embedding layer * [Feature] Add input embedding layer * [Feature] Add input embedding layer * [Feature] Add input embedding layer * [Feature] Add input embedding layer * [Feature] Add input embedding layer * [Feature] Add input embedding layer * [Feature] Add input embedding layer * [Doc] fill readme with new performance results * [Doc] Add Chinese User Guide, graph and 1.5 * [Doc] Add Chinese User Guide, graph and 1.5 * Update README.md * [Fix] Temporary remove compgcn * [Doc] Add CN user guide chapter2 * [Test] Tunning format * [Test] Tunning format * [Test] Tunning format * [Test] Tunning format * [Test] Tunning format * [Test] Section headers * [Fix] Fix format errors * [Fix] Fix format errors * [Fix] Fix format errors * [Doc] Add CN-EN EN-CN links * [Doc] Add CN-EN EN-CN links * [Doc] Copyedit chapter2 * [Doc] Copyedit chapter2 * [Doc] Remove EN in 2.1 * [Doc] Remove EN in chapter 2 * [Doc] Copyedit first 2 sections * [Doc] Copyedit first 2 sections * [Doc] copyedited chapter 2 CN * [Doc] Add chapter 3 raw texts * [Doc] Add chapter 3 preface and 3.1 * [Doc] Add chapter 3.2 and 3.3 * [Doc] Add chapter 3.2 and 3.3 * [Doc] Add chapter 3.2 and 3.3 * [Doc] Remove EN parts * [Doc] Copyediting 3.1 * [Doc] Copyediting 3.2 and 3.3 * [Doc] Proofreading 3.1 and 3.2 * [Doc] Proofreading 3.2 and 3.3 * [Doc] Add chapter 4 CN raw text. * [Clean] Remove codes in other branches * [Doc] Start to copyedit chapter 4 preface * [Doc] copyedit CN section 4.1 * [Doc] Remove EN in User Guide Chapter 4 * [Doc] Copyedit chapter 4.1 * [Doc] copyedit cn chapter 4.2, 4.3, 4.4, and 4.5. * [Doc] Fix errors in EN user guide graph feature and heterograph * [Doc] 2nd round copyediting with Murph's comments * [Doc] 3rd round copyediting with Murph's comments * [Doc] 3rd round copyediting with Murph's comments * [Doc] 3rd round copyediting with Murph's comments * [Sync] syncronize with the dgl master * [Doc] edited after Minjie's comments, 1st round * update cub Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>
46 行
2.2 KiB
ReStructuredText
46 行
2.2 KiB
ReStructuredText
.. _guide_cn-data-pipeline-savenload:
|
|
|
|
4.4 保存和加载数据
|
|
----------------------
|
|
|
|
:ref:`(English Version) <guide-data-pipeline-savenload>`
|
|
|
|
DGL建议用户实现保存和加载数据的函数,将处理后的数据缓存在本地磁盘中。
|
|
这样在多数情况下可以帮用户节省大量的数据处理时间。DGL提供了4个函数让任务变得简单。
|
|
|
|
- :func:`dgl.save_graphs` 和 :func:`dgl.load_graphs`: 保存DGLGraph对象和标签到本地磁盘和从本地磁盘读取它们。
|
|
- :func:`dgl.data.utils.save_info` 和 :func:`dgl.data.utils.load_info`: 将数据集的有用信息(python dict对象)保存到本地磁盘和从本地磁盘读取它们。
|
|
|
|
下面的示例显示了如何保存和读取图和数据集信息的列表。
|
|
|
|
.. code::
|
|
|
|
import os
|
|
from dgl import save_graphs, load_graphs
|
|
from dgl.data.utils import makedirs, save_info, load_info
|
|
|
|
def save(self):
|
|
# 保存图和标签
|
|
graph_path = os.path.join(self.save_path, self.mode + '_dgl_graph.bin')
|
|
save_graphs(graph_path, self.graphs, {'labels': self.labels})
|
|
# 在Python字典里保存其他信息
|
|
info_path = os.path.join(self.save_path, self.mode + '_info.pkl')
|
|
save_info(info_path, {'num_classes': self.num_classes})
|
|
|
|
def load(self):
|
|
# 从目录 `self.save_path` 里读取处理过的数据
|
|
graph_path = os.path.join(self.save_path, self.mode + '_dgl_graph.bin')
|
|
self.graphs, label_dict = load_graphs(graph_path)
|
|
self.labels = label_dict['labels']
|
|
info_path = os.path.join(self.save_path, self.mode + '_info.pkl')
|
|
self.num_classes = load_info(info_path)['num_classes']
|
|
|
|
def has_cache(self):
|
|
# 检查在 `self.save_path` 里是否有处理过的数据文件
|
|
graph_path = os.path.join(self.save_path, self.mode + '_dgl_graph.bin')
|
|
info_path = os.path.join(self.save_path, self.mode + '_info.pkl')
|
|
return os.path.exists(graph_path) and os.path.exists(info_path)
|
|
|
|
请注意:有些情况下不适合保存处理过的数据。例如,在内置数据集 :class:`~dgl.data.GDELTDataset` 中,
|
|
处理过的数据比较大。所以这个时候,在 ``__getitem__(idx)`` 中处理每个数据实例是更高效的方法。
|