* Use ==/!= to compare constant literals (str, bytes, int, float, tuple)
Avoid Syntax Warnings on Python >= 3.8
$ `python3`
```
>>> "" == ""
True
>>> "" is ""
<stdin>:1: SyntaxWarning: "is" with a literal. Did you mean "=="?
True
```
* Use ==/!= to compare constant literals (str, bytes, int, float, tuple)
* The start of experiments of Jiahang Li on GraphSAINT.
* a nightly build
* a nightly build
Check the basic pipeline of codes. Next to check the details of samplers , GCN layer (forward propagation) and loss (backward propagation)
* a night build
* Implement GraphSAINT with torch.dataloader
There're still some bugs with sampling in training procedure
* Test validity
Succeed in testing validity on ppi_node experiments without testing other setup.
1. Online sampling on ppi_node experiments performs perfectly.
2. Sampling speed is a bit slow because the operations on [dgl.subgraphs], next step is to improve this part by putting the conversion into parallelism
3. Figuring out why offline+online sampling method performs bad, which does not make sense
4. Doing experiments on other setup
* Implement saint with torch.dataloader
Use torch.dataloader to speed up saint sampling with experiments. Except experiments on too large dataset Amazon, we've done some experiments on other four datasets including ppi, flickr, reddit and yelp. Preliminary experimental results show consumed time and metrics reach not bad level. Next step is to employ more accurate profiler which is the line_profiler to test consumed period, and adjust num_workers to speed up sampling procedures on same certain datasets faster.
* a nightly build
* Update .gitignore
* reorganize codes
Reorganize some codes and comments.
* a nightly build
* Update .gitignore
* fix bugs
Fix bugs about why fully offline sampling and author's version don't work
* reorganize files and codes
Reorganize files and codes then do some experiments to test the performance of offline sampling and online sampling
* do some experiments and update README
* a nightly build
* a nightly build
* Update README.md
* delete unnecessary files
* Update README.md
* a nightly update
1. handle directory named 'graphsaintdata'
2. control graph shift between gpu and cpu related to large dataset ('amazon')
3. remove parameter 'train'
4. refine annotations of the sampler
5. update README.md including updating dataset info, dependencies info, etc
* a nightly update
explain config differences in TEST part
remove a sampling time variant
make 'online' an argument
change 'norm' to 'sampler'
explain parameters in README.md
* Update README.md
* a nightly build
* make online an argument
* refine README.md
* refine codes of `collate_fn` in sampler.py, in training phase only return one subgraph, no need to check if the number of subgraphs larger than 1
* Update sampler.py
check the problem on flickr is about overfitting.
* a nightly update
Fix the overfitting problem of `flickr` dataset. We need to restrict the number of subgraphs (also the number of iterations) used in each epoch of training phase. Or it might overfit when validating at the end of each epoch. The method to limit the number is a formula specified by the author.
* Set up a new flag `full` specifying if the number of subgraphs used in training phase equals to that of pre-sampled subgraphs
* Modify codes and annotations related the new flag
* Add a new parameter called `node_budget` in the base class `SAINTSampler` to compute the specific formula
* set `gpu` as a command line argument
* Update README.md
* Finish the experiments on Flickr, which is done after adding new flag `full`
* a nightly update
* use half of edges in the original graph to do sampling
* test dgl.random.choice with or without replacement with half of edges
~ next is to test what if put the calculating probability part out of __getitem__ can speed up sampling and try to implement sampling method of author
* employ cython to implement edge sampling for per edge
* employ cython to implement edge sampling for per edge
* doing experiments to test consumed time and performance
** the consumed time decreased to approximately 480s, the performance decrease about 5 points.
* deprecate cython implementation
* Revert "employ cython to implement edge sampling for per edge"
* This reverts commit 4ba4f092
* Deprecate cython implementation
* Reserve half-edges mechanism
* a nightly update
* delete unnecessary annotations
Co-authored-by: Mufei Li <mufeili1996@gmail.com>
* fix.
* fix.
* fix.
* fix.
* Fix test
* Deprecate old DistEmbedding impl, use synchronized embedding impl
* Basic imple of heterogeneous on homogenenous sampling
* make pass
* Pass C++ test
* Add python test code
* lint
* lint
* Add MultiLayerEtypeNeighborSampler
* Add unitest for single machine dataloader
* Add dist dataloader test for edge type sampler
* Fix lint
* fix
* support for per etype sample
* Fix some bug and enable distributed training with per edge sample
* fix
* Now distributed training works
* turn off some mxnet
* turn off mxnet for some dist test
* fix
* upd
* upd according to the comments
* Fix
* Fix test and now distributed works.
* upd
* upd
* Fix
* Fix bug
* remove dead code.
* upd
* Fix
* upd
* Fix
Co-authored-by: Ubuntu <ubuntu@ip-172-31-71-112.ec2.internal>
Co-authored-by: Ubuntu <ubuntu@ip-172-31-2-66.ec2.internal>
Co-authored-by: Da Zheng <zhengda1936@gmail.com>
* Create EEG-GCNN example.
* Update README.md
* Remove gitignore file.
* Update README.md
* change 'datas' to 'datasets'.
* Change train.py to main.py
* Added an entry in the indexing page.
* State "simplified version"; change how to run.
* Fix bug in contact
* Remove paper link in reference.
* Create working branch
* Add normalization of x.
* Update paper link and tags
* Update paper link in readme
* Update readme; add patient level indices
* Update readme. Add comments to models
* Update README.md
* change to with; specify location for ch and el; move note
* fix bug for note
* Add args for models; clean code.
* delete = in readme
* Add reference for spec_coh_values
* Added code for Rectifying (TypeError: unhashable type: 'slice') when copying file
* 1) added distributed preprocessing code to create ParMetis Input from CSV files
2) add code to run pm_dglpart on multiple machines
3) added support for recreating heteregenous graph from homo geneous graph based on dropped edges, as ParMetis currently only supports homogeneous graphs
* move to pandas
* Added comments and remove drop_duplicates as it was redundant
* Addressed Pr Comments
* Rename variable
* Added comment
* Added comment
* updated ReadMe
Co-authored-by: Ankit Garg <gaank@amazon.com>
Co-authored-by: Da Zheng <zhengda1936@gmail.com>
* [Model] add model example CARE-GNN
* update README
* improvements based on the review feedback
* fix missing item()
Co-authored-by: zhjwy9343 <6593865@qq.com>
* add model example GCN-based Anti-Spam
* update example index
* add usage info
* improvements as per comments
* fix image invisiable problem
* add image file
Co-authored-by: zhjwy9343 <6593865@qq.com>
* fix breakline in fakenews.py
* fix inconsistent argument name
* modify incorrect example and deprecated graph type
* modify docstring and example in knn_graph
* fix incorrect node type
Co-authored-by: Quan (Andy) Gan <coin2028@hotmail.com>
* add hilander model implementation draft
* use focal loss
* fix
* change data root
* add necessary scripts
* update download links
* update
* update example table
* fix
* update readme with numbers
* add empty folder
* only eval at the end
* set up hilander
* inform results may fluctuate
* address comments
Co-authored-by: sneakerkg <xiaotj1990327@gmail.com>
Co-authored-by: Ubuntu <ubuntu@ip-172-31-19-212.us-east-2.compute.internal>
* csr and csc creation
* fix
* fix
* fixes to adj transpose
* fine
* raise error if indptr did not match number of nodes
* fix
* huh?
* oh
Co-authored-by: Minjie Wang <wmjlyjemaine@gmail.com>