The functions `_get_consensus_stats` and
`_get_annotator_label_quality_score` take an argument
`label_quality_score_kwargs`, a dictionary of keyword arguments to pass
to `get_label_quality_scores`. When passing a
`label_quality_score_kwargs` dictionary to these functions, using the
unpacking operator is incorrect: that would be an extra level of
unpacking. The _implementations_ of these functions will unpack the
`label_quality_score_kwargs` when calling `get_label_quality_scores`.
This patch fixes the issue and adds a basic regression test.
[skip ci]
* error handling for rare classes
* change subtract to symmetric difference
* remove extra np.unique
* add warning for all instances of getting consensus labels
* error checking edits
* Add typing
Co-authored-by: Elías Snorrason <eliassno@gmail.com>
* black formatting
* update typing - labels_multiannotator will always already be converted to pd.dataframe
* make pred_probs options in typing
* add =None
* docstring edits
Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
* add basic docstring
* Changed to verbose
* Removed unessesary calculation out of get_labels_quality_m
* Update cleanlab/multiannotator.py
* comment for lost classes check so it can be grepped
* caution about setting verbose to false
* advise against verbose=false in docstring
Co-authored-by: Elías Snorrason <eliassno@gmail.com>
Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
Co-authored-by: Ulyana <ulyana@cleanlab.ai>
* fix bug in hard-coded test. generalize the test
* 🐛 cast rounded num_issues to int
np.rint outputs an array of the same shape and type as its input. num_label_issues is expected to return an integer.
Co-authored-by: Elías Snorrason <eliassno@gmail.com>
* Added base structure outline for get_ood_scores()
Added base test for get_ood_scores()
* Added warning for illogical param combo
* Addressed PR comments
* Added better unit tests
* TODO: test for correctly identifying OOD example
* Moved logic from get_ood_scores to _subtract_confident_thresholds
* TODO: Is this a cleaner way of doing this? Minimal code repeat but strange change to _subtract function
* Addressed type issue
* Switched logic for getting confident thresholds
* Fixing typecheck issues wiht labels parameter being None
* Simplified helper function. Testing type
* Fixed mypy static typing issue
* Mypy typecheck logic test
* removed uncessesary imports in util file
* typechecker debugging (add assert)
* Fixed type logic and removed confident_thresholds=None return
* Added extra arg in helper func to end of func
* Added zero-index checking for label param
* Added ood examples to outlier score notebook
* Added skeleton file structure for implementing outliers
* Make adjust_pred_probs=True by default not false
* Added base Outlier class functionality
TODO:
* test_outlier.py
* Added logic tests for function
* Updated get_ood_scores to always return confident_thresholds
* Even if none were calculated (then return is None)
* Added warning for fit that doesn't calculate confident_thresholds
* Moved get_outlier_scores and get_ood_scores to outlier.py
* Changed param dicts to dicts
* Added docstring to outlier.py
* Added proper return types
* Fixed mypy typing issues
* Switched outliers -> features; ood -> predictions naming conv
* Switched docstring to stem from fit and score functions
* Changed return of helper functions
* Fixed tutorials notebook to use OutOfDistribution class
* Moved imports to top of file
* Added option for different knn objects, addressed pr comments
* Switched params arg to init only
* Addressed PR comment for notebook, cleared notebook
* Fixed PR comments, wording.
* Testing relative links on build
* Changed ood_pred_probs scores to reflect 0 = most ood 1 = least
* Added MLP for detection outliers with pred_probs
* Fixed typos
* Added MLP Import
* Improved warning when fit call unnecessary
* Changed referenced to params dict in warnings/errors
* Added bagging+MLP classifier into notebook
* Reverted tutorial wording
* OOD tutorial improvements
* cleanup OOD documentation
* adopting->using
Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
* 🎨 remove wildcard imports
* 🏷️ review type annotations in display_issues
Restrict nested lists with `List`. Keep unrestricted lists as `list`. Remove hard-coded tags from docstrings, should be auto-tagged in later PR.
* ♻️ use `isinstance` for type-check at runtime
* 🏷️ review type annotations in common_label_issues and filter_by_token
Remove hard-coded tags in docstrings. Should be auto-tagged in later PR.
* ✅ improve test coverage of display issues
No part of the function is easily testable except ensuring it completes exection. Some parts handle edge cases that were never reached during testing.
* ✅ parametrize tests for coverage on common_label_issues and filter_by_token
* ✨ search tokenized sentence for coloring
Searches through the list of tokens before trying to match substrings. Thanks for this code suggestion Eric!
* 🚧 fix signature in all calls to color_sentence
* ✅ update color_sentence test after changing its api
(sentence, word) -> (word, tokens)
* ⏪ use regex for coloring tokens in sentence
Find word boundaries with regex, use builtin replace() as fallback w/o boundaries.
Closes#288
* 🚑 invert fallback condition
Use replace if NO substitutions were made with regex.
* ✅ add test fixture for get_label_quality_scores
* ✅ test softmin_sentence_score
include test cases for temperature limits
* ♻️ cleanup softmin_sentence_score
Remove unused keyword-only args, Fix tag in docstring, Change format of nested functions.
* ⚡ specialize edge-case temperature=inf in softmin_sentence_score
* ♻️ simplify temperature lookup
* ♻️ cleanup get_label_quality_scores
Remove unused args/variables. Update parameter list in docstring. Rename parameter of inner function. Function always returns a tuple.
* 🏷️ tag token_scores as optional
* ✅ test raised error
* 🩹 skip untestable elif statement
the elif statement only gets partial coverage because it can't evaluate to False due to the `assert sentence_score_method` at the start of the function
* 🏷️ restrict parameters for list types
* 🐛 only process characters in input token
Example: process_token("Cleanlab", [("C", "a"), ("a", "C")]) should return "aleCnlCb", not "CleCnlCb".
* ✅ use all sentences in test_get_sentence
* ✅ add test cases to test_filter_sentence
* ✅ extent test cases in test_mapping
* 📝 clean up docstrings
Restrict arg types based on docstrings. Fix punctuation and typos. Add examples to docstring.
* ✅ split tests for filter_sentence
* ✅ extend test_merge_probs
* ✅ test merge_probs with ignored/normalized columns in probs
* ✅ extend test cases for get_sentences
* ⚰️ remove unused pandas import
* 🏷️ pass strict mypy check
We ignore np.max as it is untyped.
No issues found in token_classification_utils.py by running
```
mypy --install-types --non-interactive --strict cleanlab/internal/token_classification_utils.py
```
* 👷 add strict type-checking in CI
* 💚 use strict type-check for single file
* ⏪ remove strict type check in CI
* refactor: 🏷️ use np.ndarray type instead of npt.NDArray
* ⏪ go back to generic np.ndarray type
* ♻️ always return tuple in filter_sentence
Remove unused argument+docstring. Simplify relevant unit tests.
* 🔥 resolve comments on typing
Remove ignore-comments. Remove duplicate tag in docstring. Remove unused imports.
* 🔥 remove duplicate tag in docstring
* validation_func docstring
* torch,tf compatibility+tests
* keras test
* skip tests if python < 3.7
* pytorch numpy int bug on windows
* make tensorflow test work on windows
* move tf env variable setting
* pytorch test increase epochs
* install cpu-tensorflow on windows CI
* torch test optimizer to adam
* fix bugs in shuffled TF dataset
* dummy unit test for TF on windows
* dummy code for TF windows testing
* deal with np.int bug on windows
* remove windows debugging code
* docstrings for new functionality
* address merge conflicts
* reformat after merge
* addressed comments
* 🏷️ annotate function args and return values
Starting with the validation module:
- I think (X, y) might need some custom Union type to handle both numpy arrays and pandas dataframe, etc.
- All of the "assert" functions return None.
Ref #307
* refactor: 🏷️ swap npt.NDArray -> np.ndarray
np.ndarray seems more consistent with the rest of the repo.
Maybe it's necessary to go back to npt.NDArray when disallowing generics?
See numpy docs: https://numpy.org/devdocs/reference/typing.html#numpy.typing.NDArray
* refactor: 🔥 remove unused import
* 🏷️ unconstrain X input types
* 🐛 handle label type issues
- Have to restrict the output of labels_to_array to pass mypy checks.
- Returning the values of pd.Series isn't type-stable.
* 🏷️ include np.generic in arg-type union
* test: ✅ test labels_to_array
* 🏷️ add type aliases for X and y
* 🚨 ignore type-checks for pandas indexing assertions
CI typechecker runs on Python 3.10 which gives this error:
'cleanlab/internal/validation.py:125: error: No overload variant of "__getitem__" of "_iLocIndexerSeries" matches argument type "List[int]"'
It should be fine to let mypy ignore these expressions as they don't return anything.
* 🥅 specify errors to ignore
"type: ignore" doesn't pass strict mypy type-checks unless the specific errors are provided
* 🏷️ annotate label series to array
* 🗑️ change labels arg -> y in CleanLearning.fit()
Set `label` as an optional keyword-only argument.
Anyone still using it in this method should get a deprecation warning.
Fixes#281
* 📝 add note for y/labels in docstring
* 🥅 make y an optional positional arg.
Should now resolve deprecated signatures.
* 📝 labels -> y in module docstring
* ⏪ revert "label -> y deprecation"
This reverts commit ab319a0cca2cec715a84eb5f628bbab7706c5f9c.
This reverts commit 1b739002d848e1f0acb6390a666f6e695e25fcaa.
This reverts commit 88bb6c3bcca298dab414c3cb20101783d78d35d1.
This reverts commit d988e3c3932107e779598d02d8f16d7e6671e9e7.
* ✨ add y alias for labels
Resolves#281
* Added support to build KNN graph with only training data
- Added default classifier for get_knn_distance_ood_scores()
* Added unit test for checking default KNN model is used when nbrs=None
* Changed default num neighbors from 10 to K.
* Changed unit test to check score sums
* Fixed k=15
* Added support to build KNN graph with only training data
- Added default classifier for get_knn_distance_ood_scores()
- Added unit test for checking default KNN model is used when nbrs=None
- Changed default num neighbors from 10 to K.
- Changed unit test to check score sums
* Improved code readibility, extended unit test to check for default k
* Changed naming convention from get_knn_distance_ood_scores to get_outlier_scores
- Changed nbrs to knn
- Improved unit test readability
* Made unit test more robust to check if user-set value is passed
* changed classifier -> estimator
* fixed test warning handle
* Updated function headers to better definition
* Improved header writing added functionality to avoid training/testing with identical features
* Updated docstring, updated handling of features=None and test for it
* Added ValueError for k>len(features) and test to catch ValueError
* Updated argument types to Optional, added runtime typecheck.
* Added unit test to check TypeError
* Changed Exception type thrown when knn=None, features=None
* TypeError to ValueError
* Reversed scoring for outlier severity
* Added test to check t parameter
* Added t parameter for global rescaler
* Improved function definition concerning 't' parameter
* Fixed tests
* Removed repeated calls
* Added assertion to make sure X_ood is always smallest outlier score
* Improved logic for checking X_ood has the smallest score