* Added runnable package versioning for tutorials
Added quickstart at top of tutorials
* Addressed PR comments
* Fixed quickstart to include cleanlearning
* Set correct requirements.txt
* Updated y to labels
* Fixed y-> labels labeling issues
* Updated quickstart message/removed extra dependencies
* remove keras from package-versions
* more concise
* remove pathlib as explicit requirement
* remove pathlib version
* rewording
* true label -> given label
* remove venv from kernelspec
Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
* cleanlearning tips
* addessed comments
* add a few general tips at the end
* minor edits
* format pred_probs as variable type
* bug fix to pass lint (missing comma)
* fix format for linter
Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>
* Updating tutorials hyperlink to 2.0.0 release
* revert back to v.2.0.0 links instead of stable
Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
* Added outlier detection tutorial into docs
* Switched to pytorch model/dataset implementation
* Fixed image normalization
* Fixed outlier detection for test set
* Added outlier thresholding into tutorial.
* Cleaned output
* Removed unused imports and renamed notebook
* Added outliers notebook to PR
* Added quickstart into tutorial
* Best quickstart header
* Cleaned cell output
* Changed to use subset of original data for speed
* Fixed randomness
* Cleared outputs
* fixed metadata tags
* Fixed metadata
* Cleaned kernel and verified output
* Improved unit test
Changed labels references to classes where apropriate
The previous link took you to a tutorials page with lots of options including tutorials that were for data-centric AI workflows, and computing cross validated probs, etc. if this is supposed to be the very first thing users click on to get started, we want it to take them straight to a place where they can get started in 5 minutes, not a list of things where they have to figure out what to click next.
Updated to fix this, which also reduced the length :)
* update N in classification.py
* minor docstring changes on K-1 classes
* use K in shape
* minor grammar fixes
* docs language improvements
Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
* Added support to build KNN graph with only training data
- Added default classifier for get_knn_distance_ood_scores()
* Added unit test for checking default KNN model is used when nbrs=None
* Changed default num neighbors from 10 to K.
* Changed unit test to check score sums
* Fixed k=15
* Added support to build KNN graph with only training data
- Added default classifier for get_knn_distance_ood_scores()
- Added unit test for checking default KNN model is used when nbrs=None
- Changed default num neighbors from 10 to K.
- Changed unit test to check score sums
* Improved code readibility, extended unit test to check for default k
* Changed naming convention from get_knn_distance_ood_scores to get_outlier_scores
- Changed nbrs to knn
- Improved unit test readability
* Made unit test more robust to check if user-set value is passed
* changed classifier -> estimator
* fixed test warning handle
* Updated function headers to better definition
* Improved header writing added functionality to avoid training/testing with identical features
* Updated docstring, updated handling of features=None and test for it
* Added ValueError for k>len(features) and test to catch ValueError
* Updated argument types to Optional, added runtime typecheck.
* Added unit test to check TypeError
* Changed Exception type thrown when knn=None, features=None
* TypeError to ValueError
* Reversed scoring for outlier severity
* Added test to check t parameter
* Added t parameter for global rescaler
* Improved function definition concerning 't' parameter
* Fixed tests
* Removed repeated calls
* Added assertion to make sure X_ood is always smallest outlier score
* Improved logic for checking X_ood has the smallest score
Our package doesn't have type annotations everywhere, so we can't use
mypy in strict mode just yet. Still, adding type checking in CI is
valuable, so we don't have unchecked annotations in our code.
This patch includes basic fixes to make type checking pass, including
switching the incorrect `np.array` type annotation for `np.ndarray` and
adding some assertions for flow-sensitive typing.
* Create FAQ page in cleanlab docs
* Add FAQ notebook to store answers to frequently asked questions
* Improved formatting issues for faq page
- added faq to sidebar
* Improved formatting issues for faq page
- added faq to sidebar
- changed email to comply with CLA
* Removed faq numbering
- Removed notebook output metadata
- Added notebook to index for left table link
* Text and example rewritten for clarity
- How do I format labels for cleanlab-- example improved
- How do I format labels for cleanlab-- text amended
- Can't find an answer to your question-- added links
* add links to issues/slack
Co-authored-by: Ulyana Tkachenko <uly@ulyana.lan>
Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
* cleanlearning w dfs
* work for sparse matrix as well
* simplify logic of labels_to_array and extend types
* address pr feedback
* add unit test
* rare label dataframe
* modularize subsetting code
* series rarelabel test
* replace cal.com with slack/email
* Add general method to find num_classes from labels
* compute num_classes with pred_probs.shape[1]
* fix broken commits, address 2nd round of comments
Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>
Add clarification of the labels format requirements for all major API functions.
* Fix broken link
* Clarify reqs for labels format. rank.py does not yet support multi_label
* add double ticks to code in docstrings
* clarify docstring
* further clarify multi_label vs single label labels reqs
* also add new docstring to confident joint
* example label formatting
* example label format in text tutorial
* example label format for tabular tutorial
* y -> full_labels
* example label format for audio tutorial
* clarify input shapes in image tutorial
* clarify input shapes in text tutorial
* clarify input shapes in tabular tutorial
* clarify input shapes for audio tutorial
* Add KNN distance OOD scoring function and unit tests
* Update KNN distance OOD scoring function
* Change query_features to features in unit tests for KNN distance OOD scoring function
* Update KNN distance OOD scoring function
* Update tests for KNN distance OOD scoring function to use auto for algo
* Allow k=None for KNN Distance OOD score
* added class to use cleanlab with tensorflow and huggingface models
* - Documentation refactoring
- predict and predict_proba functions work on new test data too
* Add log_loss_search weighting method for ensemble label quality scoring function
* Update log_loss_search weighting method
* Add test for log_loss_search method
* Add parameter to Ensemble label quality scoring function for t values in log_loss_search method
* Update ensemble label quality scoring function docstring
* Update ensemble label quality scoring function comments
* Update docstring in ensemble label quality scoring function
* Update docstring in ensemble label quality scoring function
* Update docstring in ensemble label quality scoring function
* Modify verbose printout for log_loss_search
* Add clipping of pred_prob when calculating weights for log_loss_search
* Add clipping of pred_prob and renormalization when calculating weights for log_loss_search
* Add comments for log_loss_search weighting scheme
The issue that was introduced in coverage 6.3 has been fixed:
https://github.com/nedbat/coveragepy/issues/1310#issuecomment-1129894701.
We can't just upgrade to `coverage` or `coverage>=6.4` because the
former could install bad versions of coverage (e.g. 6.3), and the latter
is unsupported on Python 3.6, which we want to continue supporting.
This patch just bans coverage 6.3 / 6.3.x, so with Python 3.7+, we'll
use the latest version of coverage, and with Python 3.6, we'll use the
latest supported version of coverage that's not a 6.3 release.