文件历史

177 次代码提交

作者 SHA1 备注 提交日期
Ulyana 5da046da9e Added support for returning ranked issue idxs (#459)
* Added support for returning ranked issue idxs
- code
- tests
- docstring

Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>
2022-09-16 20:20:52 -07:00
Hui Wen 06ed233a25 Error handling for rare classes (#455)
* error handling for rare classes

* change subtract to symmetric difference

* remove extra np.unique

* add warning for all instances of getting consensus labels

* error checking edits

* Add typing

Co-authored-by: Elías Snorrason <eliassno@gmail.com>

* black formatting

* update typing - labels_multiannotator will always already be converted to pd.dataframe

* make pred_probs options in typing

* add =None

* docstring edits

Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>

* add basic docstring

* Changed to verbose

* Removed unessesary calculation out of get_labels_quality_m

* Update cleanlab/multiannotator.py

* comment for lost classes check so it can be grepped

* caution about setting verbose to false

* advise against verbose=false in docstring

Co-authored-by: Elías Snorrason <eliassno@gmail.com>
Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
Co-authored-by: Ulyana <ulyana@cleanlab.ai>
2022-09-16 19:26:29 -07:00
Jonas Mueller 7256cd3546 More improvements to token classification code and documentation (#452)
* improved token docs/code

* format paragraphs in docstrings

* fix typo

Co-authored-by: Elías Snorrason <eliassno@gmail.com>
2022-09-16 14:04:45 +00:00
Elías Snorrason 95b0742342 Make softmin_sentence_score a private function (#449)
This scoring function is only used internally. Might as well be private.
2022-09-15 23:19:15 +00:00
Curtis G. Northcutt 1b813d9666 fix bug in hard-coded test. generalize the test (#448)
* fix bug in hard-coded test. generalize the test

* 🐛 cast rounded num_issues to int

np.rint outputs an array of the same shape and type as its input. num_label_issues is expected to return an integer.

Co-authored-by: Elías Snorrason <eliassno@gmail.com>
2022-09-15 18:16:56 -04:00
Ulyana 61f650e357 Merge pull request #426 from ulya-tkch/ulya-outlier-array-like
Updated labels to allow array_like
2022-09-14 21:21:33 -07:00
Hui Wen 5603b1c8a6 Extend KerasWrapper to Functional API (#434)
* add keras functional api wrapper

* add breif docstring

* minor unittest edit

* edit unittest

* update docstring

* make unittest less strict
2022-09-14 13:10:12 -07:00
Ulyana 13d6cfc9ce Used internal function for typechecking labels 2022-09-12 16:54:07 -07:00
Elías Snorrason 9a109182b1 Update argument given_labels -> labels in token classification module (#423)
* given_labels -> labels

only applies to the argument to `display_issues`

Fixes #418

* update kwargs in test for display_issues

* apply black formatter
2022-09-11 12:38:38 -07:00
Ulyana a6fb96f082 Add test 2022-09-09 16:58:42 -07:00
Ulyana 9dfa001068 Implementing get_ood_scores function (#338)
* Added base structure outline for get_ood_scores()
Added base test for get_ood_scores()

* Added warning for illogical param combo

* Addressed PR comments

* Added better unit tests
* TODO: test for correctly identifying OOD example

* Moved logic from get_ood_scores to _subtract_confident_thresholds

* TODO: Is this a cleaner way of doing this? Minimal code repeat but strange change to _subtract function
* Addressed type issue

* Switched logic for getting confident thresholds

* Fixing typecheck issues wiht labels parameter being None

* Simplified helper function. Testing type

* Fixed mypy static typing issue

* Mypy typecheck logic test

* removed uncessesary imports in util file

* typechecker debugging (add assert)

* Fixed type logic and removed confident_thresholds=None return

* Added extra arg in helper func to end of func

* Added zero-index checking for label param

* Added ood examples to outlier score notebook

* Added skeleton file structure for implementing outliers

* Make adjust_pred_probs=True by default not false

* Added base Outlier class functionality

TODO:
* test_outlier.py

* Added logic tests for function

* Updated get_ood_scores to always return confident_thresholds
* Even if none were calculated (then return is None)

* Added warning for fit that doesn't calculate confident_thresholds

* Moved get_outlier_scores and get_ood_scores to outlier.py

* Changed param dicts to dicts

* Added docstring to outlier.py

* Added proper return types

* Fixed mypy typing issues

* Switched outliers -> features; ood -> predictions naming conv

* Switched docstring to stem from fit and score functions

* Changed return of helper functions

* Fixed tutorials notebook to use OutOfDistribution class

* Moved imports to top of file

* Added option for different knn objects, addressed pr comments

* Switched params arg to init only

* Addressed PR comment for notebook, cleared notebook

* Fixed PR comments, wording.

* Testing relative links on build

* Changed ood_pred_probs scores to reflect 0 = most ood 1 = least

* Added MLP for detection outliers with pred_probs

* Fixed typos

* Added MLP Import

* Improved warning when fit call unnecessary

* Changed referenced to params dict in warnings/errors

* Added bagging+MLP classifier into notebook

* Reverted tutorial wording

* OOD tutorial improvements

* cleanup OOD documentation

* adopting->using

Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
2022-09-07 10:52:55 -07:00
Elías Snorrason 8371edffc9 🐛 escape special regex characters (#404)
Fixes #403
2022-09-06 09:28:59 -07:00
Elías Snorrason 9f030a65ac Cleanup-summary (#396)
* 🎨 remove wildcard imports

* 🏷️ review type annotations in display_issues

Restrict nested lists with `List`. Keep unrestricted lists as `list`. Remove hard-coded tags from docstrings, should be auto-tagged in later PR.

* ♻️ use `isinstance` for type-check at runtime

* 🏷️ review type annotations in common_label_issues and filter_by_token

Remove hard-coded tags in docstrings. Should be auto-tagged in later PR.

*  improve test coverage of display issues

No part of the function is easily testable except ensuring it completes exection. Some parts handle edge cases that were never reached during testing.

*  parametrize tests for coverage on common_label_issues and filter_by_token
2022-09-05 17:21:26 -07:00
Elías Snorrason 71521e90d3 Match token/s in color_sentence (#397)
*  search tokenized sentence for coloring

Searches through the list of tokens before trying to match substrings. Thanks for this code suggestion Eric!

* 🚧 fix signature in all calls to color_sentence

*  update color_sentence test after changing its api

(sentence, word) -> (word, tokens)

*  use regex for coloring tokens in sentence

Find word boundaries with regex, use builtin replace() as fallback w/o boundaries.

Closes #288

* 🚑 invert fallback condition

Use replace if NO substitutions were made with regex.
2022-09-05 16:30:29 -07:00
Elías Snorrason afa217cb4a Fix typing for find_label_issues (#391) 2022-09-03 12:30:07 -07:00
Elías Snorrason c884dc5171 Cleanup token_classification.rank module (#393)
*  add test fixture for get_label_quality_scores

*  test softmin_sentence_score

include test cases for temperature limits

* ♻️ cleanup softmin_sentence_score

Remove unused keyword-only args, Fix tag in docstring, Change format of nested functions.

*  specialize edge-case temperature=inf in softmin_sentence_score

* ♻️ simplify temperature lookup

* ♻️ cleanup get_label_quality_scores

Remove unused args/variables. Update parameter list in docstring. Rename parameter of inner function. Function always returns a tuple.

* 🏷️ tag token_scores as optional

*  test raised error

* 🩹 skip untestable elif statement

the elif statement only gets partial coverage because it can't evaluate to False due to the `assert sentence_score_method` at the start of the function
2022-09-02 18:26:07 -07:00
Elías Snorrason 6ec5b173dd Cleanup token classification utils (#390)
* 🏷️ restrict parameters for list types

* 🐛 only process characters in input token

Example: process_token("Cleanlab", [("C", "a"), ("a", "C")]) should return "aleCnlCb", not "CleCnlCb".

*  use all sentences in test_get_sentence

*  add test cases to test_filter_sentence

*  extent test cases in test_mapping

* 📝 clean up docstrings

Restrict arg types based on docstrings. Fix punctuation and typos. Add examples to docstring.

*  split tests for filter_sentence

*  extend test_merge_probs

*  test merge_probs with ignored/normalized columns in probs

*  extend test cases for get_sentences

* ⚰️ remove unused pandas import

* 🏷️ pass strict mypy check

We ignore np.max as it is untyped.

No issues found in token_classification_utils.py by running
```
mypy --install-types --non-interactive --strict cleanlab/internal/token_classification_utils.py
```

* 👷 add strict type-checking in CI

* 💚 use strict type-check for single file

*  remove strict type check in CI

* refactor: 🏷️ use np.ndarray type instead of npt.NDArray

*  go back to generic np.ndarray type

* ♻️ always return tuple in filter_sentence

Remove unused argument+docstring. Simplify relevant unit tests.

* 🔥 resolve comments on typing

Remove ignore-comments. Remove duplicate tag in docstring. Remove unused imports.

* 🔥 remove duplicate tag in docstring
2022-09-01 17:07:39 -07:00
Hui Wen 960c2b4ac1 CL functionality for multiannotator data (#333)
* setup multiannotator functionality

* add docstring

* add unittests

* edit get label quality with nan helper func

* docstring edits

* handle non overlapping annotators in get_annotator_agreement_with_annotators

* return NaN annotator quality scores for non-overlapping annotators

* add unittest

* add unittests for get_consensus_label tiebreaks

* np.array -> np.ndarray

* add new annotator quality method

* docstring edits

* address comments, docstring changes

* add method to compute improved consensus labels

* separate detailed_label_quality return

* elif statement typo

* changed get_majority_vote_label

* add 'best_quality' as a consensus method

* change lqs kwargs naming

* change lqs kwargs naming

* change unittest function names

* docstring edits

* change get_worst_class to return class with lowest agreement with consensus

* edit method to get annotator lqs, change quality_of_consensus -> consensus_quality_score

* remove unnecessary args

* satisfy mypy checks

* fix mypy issue

* add unittests

* reshuffled order of functions

* address comments

* docstring edit

* change return to always return dict

* add clipping to prevent negative pred_probs

* add tutorial

* add multiannotator to tutorial index

* add newline

* add tiebreaks for best_quality labels and worst_class

* change auto method name to crowdlab

* fix indexing for non numeric anno names

* add test for nonnumeric anno names

* update unittest

* address comments [partially complete]

* edit metadata

* missing /details tag

* bugfix to allow pd Int64 types

* update tutorial

* address comments
2022-08-30 11:29:06 -07:00
Eric Wang 1bad2f82a2 Adding functionality for cleanlab to find label errors in token classification datasets (#347)
* add token_classification functionality

* add typing

* fix typing

* add typing

* fixed typing and black

* fix typing
2022-08-30 09:17:07 -07:00
Jonas Mueller b93fdebf30 Add compatibility for tensorflow and pytorch Dataset objects (#311)
* validation_func docstring

* torch,tf compatibility+tests

* keras test

* skip tests if python < 3.7

* pytorch numpy int bug on windows

* make tensorflow test work on windows

* move tf env variable setting

* pytorch test increase epochs

* install cpu-tensorflow on windows CI

* torch test optimizer to adam

* fix bugs in shuffled TF dataset

* dummy unit test for TF on windows

* dummy code for TF windows testing

* deal with np.int bug on windows

* remove windows debugging code

* docstrings for new functionality

* address merge conflicts

* reformat after merge

* addressed comments
2022-07-27 21:42:13 -07:00
Elías Snorrason 8b60a381e4 validation.py: Annotate function args and return values (#317)
* 🏷️ annotate function args and return values

Starting with the validation module:

- I think (X, y) might need some custom Union type to handle both numpy arrays and pandas dataframe, etc.
- All of the "assert" functions return None.

Ref #307

* refactor: 🏷️ swap npt.NDArray -> np.ndarray

np.ndarray seems more consistent with the rest of the repo.
Maybe it's necessary to go back to npt.NDArray when disallowing generics?

See numpy docs: https://numpy.org/devdocs/reference/typing.html#numpy.typing.NDArray

* refactor: 🔥 remove unused import

* 🏷️ unconstrain X input types

* 🐛 handle label type issues

- Have to restrict the output of labels_to_array to pass mypy checks.
- Returning the values of pd.Series isn't type-stable.

* 🏷️ include np.generic in arg-type union

* test:  test labels_to_array

* 🏷️ add type aliases for X and y

* 🚨 ignore type-checks for pandas indexing assertions

CI typechecker runs on Python 3.10 which gives this error:

'cleanlab/internal/validation.py:125: error: No overload variant of "__getitem__" of "_iLocIndexerSeries" matches argument type "List[int]"'

It should be fine to let mypy ignore these expressions as they don't return anything.

* 🥅 specify errors to ignore

"type: ignore" doesn't pass strict mypy type-checks unless the specific errors are provided

* 🏷️ annotate label series to array
2022-07-26 12:58:37 -07:00
Elías Snorrason cecaf7aec6 Add y argument as alternative to labels in CleanLearning.fit() (#322)
* 🗑️ change labels arg -> y in CleanLearning.fit()

Set `label` as an optional keyword-only argument.
Anyone still using it in this method should get a deprecation warning.

Fixes #281

* 📝 add note for y/labels in docstring

* 🥅 make y an optional positional arg.

Should now resolve deprecated signatures.

* 📝 labels -> y in module docstring

*  revert "label -> y deprecation"

This reverts commit ab319a0cca2cec715a84eb5f628bbab7706c5f9c.
This reverts commit 1b739002d848e1f0acb6390a666f6e695e25fcaa.
This reverts commit 88bb6c3bcca298dab414c3cb20101783d78d35d1.
This reverts commit d988e3c3932107e779598d02d8f16d7e6671e9e7.

*  add y alias for labels

Resolves #281
2022-07-26 00:59:10 -07:00
Jonas Mueller a0e1227ae5 Allow KNN object to be returned by get_outlier_scores, Improved Tutorial (#319)
* knn return, improved tutorial

* typing fix for tuple

* add timm to docs-requirements

* language improvements

* declare typing of outlier_scores
2022-07-25 12:40:58 -07:00
Ulyana fef76ba357 Added support to build KNN graph with only training data (#305)
* Added support to build KNN graph with only training data
- Added default classifier for get_knn_distance_ood_scores()

* Added unit test for checking default KNN model is used when nbrs=None

* Changed default num neighbors from 10 to K.

* Changed unit test to check score sums

* Fixed k=15

* Added support to build KNN graph with only training data

- Added default classifier for get_knn_distance_ood_scores()

- Added unit test for checking default KNN model is used when nbrs=None

- Changed default num neighbors from 10 to K.

- Changed unit test to check score sums

* Improved code readibility, extended unit test to check for default k

* Changed naming convention from get_knn_distance_ood_scores to get_outlier_scores

- Changed nbrs to knn
- Improved unit test readability

* Made unit test more robust to check if user-set value is passed

* changed classifier -> estimator
* fixed test warning handle

* Updated function headers to better definition

* Improved header writing added functionality to avoid training/testing with identical features

* Updated docstring, updated handling of features=None and test for it

* Added ValueError for k>len(features) and test to catch ValueError

* Updated argument types to Optional, added runtime typecheck.

* Added unit test to check TypeError

* Changed Exception type thrown when knn=None, features=None

* TypeError to ValueError

* Reversed scoring for outlier severity
* Added test to check t parameter
* Added t parameter for global rescaler

* Improved function definition concerning 't' parameter

* Fixed tests
* Removed repeated calls
* Added assertion to make sure X_ood is always smallest outlier score

* Improved logic for checking X_ood has the smallest score
2022-07-14 10:36:10 -07:00
Jonas Mueller 9c543c6cba error for missing classes, consistency on determining num_classes, code cleanup (#308)
* edge cases

* unique_classes in get_confident_thresholds
2022-07-12 14:45:20 -07:00
Jonas Mueller 0163751192 Proper validation of labels values/format across package (#301)
* better checks and code

* O(1) pred_probs check, suppress cleanlearning print

* fix verbose docstring default
2022-07-01 23:53:30 -07:00
Hui Wen f62bc36e54 Allow CleanLearning to use validation data in each fold (#295)
* allow CleanLearning to use val data in each fold

* add unittest for using val data in CleanLearning
2022-06-28 22:04:22 -07:00
Jonas Mueller ffd6fc1b35 Make CleanLearning work with pandas and other non-numpy feature objects X (#285)
* cleanlearning w dfs

* work for sparse matrix as well

* simplify logic of labels_to_array and extend types

* address pr feedback

* add unit test

* rare label dataframe

* modularize subsetting code

* series rarelabel test

* replace cal.com with slack/email

* Add general method to find num_classes from labels

* compute num_classes with pred_probs.shape[1]

* fix broken commits, address 2nd round of comments

Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>
2022-06-24 04:01:55 -07:00
Johnson Kuan d902197191 Add KNN distance OOD scoring function and unit tests (#268)
* Add KNN distance OOD scoring function and unit tests

* Update KNN distance OOD scoring function

* Change query_features to features in unit tests for KNN distance OOD scoring function

* Update KNN distance OOD scoring function

* Update tests for KNN distance OOD scoring function to use auto for algo

* Allow k=None for KNN Distance OOD score
2022-06-01 17:14:32 -07:00
Johnson Kuan 04877081f3 Add Negative Log Loss Weighting Scheme for Ensemble Label Quality Score (#267)
* Add log_loss_search weighting method for ensemble label quality scoring function

* Update log_loss_search weighting method

* Add test for log_loss_search method

* Add parameter to Ensemble label quality scoring function for t values in log_loss_search method

* Update ensemble label quality scoring function docstring

* Update ensemble label quality scoring function comments

* Update docstring in ensemble label quality scoring function

* Update docstring in ensemble label quality scoring function

* Update docstring in ensemble label quality scoring function

* Modify verbose printout for log_loss_search

* Add clipping of pred_prob when calculating weights for log_loss_search

* Add clipping of pred_prob and renormalization when calculating weights for log_loss_search

* Add comments for log_loss_search weighting scheme
2022-05-26 16:55:04 -07:00
Johnson Kuan a1f8be034b Allow users to pass custom weights for ensemble label quality scoring (#255)
* Allow user to pass custom_weights to ensemble scoring method

* Add tests for ensemble scoring with custom_weights

* Add check to make sure length of custom_weights matches len(pred_probs_list)

* Update tests for usage of custom_weights in ensemble scoring
2022-05-10 18:00:28 -07:00
Jonas Mueller 74c6d5b98c unittest formatting 2022-05-04 02:52:18 -07:00
Jonas Mueller 7ed44c02ad fix sample_weight unit-test 2022-05-04 02:45:15 -07:00
Jonas Mueller 2c5a0a1f3b test sample_weight 2022-05-04 02:37:23 -07:00
Curtis G. Northcutt fac6c112aa health_summary bug fixes + notebook support (#212)
* bug fixes and jupyter notebook support added

* Increase test coverage.

* Fix bug. will always try to display if possible.
2022-04-15 14:54:00 -04:00
Jonas Mueller d1a4bc86fd Returns DataFrame type from CleanLearning functions (#199)
* df return type, need tests still

* Add pandas as a dependency

We already decided that pandas will be a dependency of cleanlab (also
used in the dataset module, see
https://github.com/cleanlab/cleanlab/pull/182).

* Tweak documentation

* addressed comments

* remove lazy import

* address 2nd round comments

* unit tests

* improve codecov

* Fix typo

* methods to save more space

* nocover statements for prints

* extra nocover

* nocover warnings

* test docstring formatting

* test docstring formatting2

* test docstring formatting2

* move compress to helper, find-label docs params

* readded stuff lost in merge conflict

* addressed remaining PR review comments

* docs formatting

* docs formatting2

* docs formatting3

* docs formatting4

* docs formatting5

* docs formatting5

* docs formatting6

* docs formatting7

* docs formatting8

* docs formatting9

* docs formatting19

* docs formatting20

* docs formatting20

* docs formatting21

* code formatting

* fix a bug where confident joint isnt computed

The confident joint wasn't getting computed if noise_matrix was passed in and pred_probs was not passed in. But that's bad because it stops workflows like:

```python
cl = CleanLearning()
cl.fit(data, labels, noise_matrix=noise_matrix)
cleanlab.dataset.health_summary(labels, confident_joint=cl.confident_joint)
```

* fixed bug from last commit. code in wrong place.

* print overwrite bugfix

Co-authored-by: Anish Athalye <me@anishathalye.com>
Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>
2022-04-13 16:00:41 -04:00
Curtis G. Northcutt c1d27cf0bf Added thorough testing. ready for release. 2022-04-13 00:39:57 -04:00
Anish Athalye eae292458f Merge branch 'master' into dataset_module 2022-04-11 12:45:21 -04:00
Curtis G. Northcutt 68d9c75e93 Add initial test for dataset module. 2022-04-10 14:38:10 +00:00
Anish Athalye 0dc384a4ed Merge branch 'docs-cleanup' 2022-04-09 15:16:09 -04:00
Anish Athalye 63393237d6 Revise noise_generation 2022-04-09 09:25:56 -04:00
Anish Athalye 8770e977d3 Remove deprecated functions
No reason to have any deprecated functions in a backwards-incompatible
release like 2.0.
2022-04-09 09:01:08 -04:00
Anish Athalye 0be3c70a6e Move noise_generation into benchmarking module 2022-04-08 20:08:19 -04:00
Anish Athalye 947ea9e34a Fix typo in test (#186)
This test had a typo in it ("pred_neq_given" rather than
"predicted_neq_given"), but the way the test was written, it didn't
fail, because GridSearchCV raises a warning if some fits fail. This
patch fixes the typo and adds the assertion that no warnings are raised
during the grid search.
2022-04-07 12:12:50 -07:00
Anish Athalye e76f2923c0 Remove deprecated test
It is annoying to re-enable this test because it's annoying to install
fastText / it doesn't run on all platforms we test in CI. Furthermore,
the data isn't available anymore (at least at the URL specified in
`get_cooking_stackexchange_data.sh`).
2022-04-07 06:21:16 -04:00
Jonas Mueller faac915740 mv example_models -> experimental 2022-04-07 01:25:03 -07:00
Jonas Mueller 423e5b0a07 Polish the APIs and file-structure to prepare for 2.0 release (#181)
* Makes some methods private that are not intended to be user-facing.
* Adds experimental module with fasttext.py and coteaching.py
* Adds header descriptions to code files which will render in docs
* Many miscellaneous fixes
2022-04-06 21:04:06 -07:00
Curtis G. Northcutt 2de67cbc34 Simple fix to Issue 158 (and potentially other issues) (#178)
* CleanLearning = Machine Learning with cleaned data

* Replace lnl instance naming with cl everywhere (CleanLearning)

* Add support for multi-class as well

* Add test based on #158

Co-authored-by: Anish Athalye <me@anishathalye.com>
2022-04-06 18:55:18 -04:00
Curtis G. Northcutt c4e84624e9 CleanLearning = Machine Learning with cleaned data (#177)
* CleanLearning = Machine Learning with cleaned data

* Replace lnl instance naming with cl everywhere (CleanLearning)

* replace rp (rank pruning) with cl (clearn learning) everywhere

* Clarifying comments. remove unnecessary newlines. fix spelling err
2022-04-06 17:33:04 -04:00
Jonas Mueller d3eb08e75a added LearningWithNoisyLabels.find_label_issues instance method (#157)
* added LearningWithNoisyLabels.find_label_issues instance method

* LNL.find_label_issues no longer memoizes

* verbose unit test coverage

* Fixed all issues in PR. Fixed confident joint usage in LNL.find_label_issues. Fixed other minor bugs. Added tests

Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>
Co-authored-by: Anish Athalye <me@anishathalye.com>
2022-04-06 12:24:27 -04:00