* ✅ add test fixture for get_label_quality_scores
* ✅ test softmin_sentence_score
include test cases for temperature limits
* ♻️ cleanup softmin_sentence_score
Remove unused keyword-only args, Fix tag in docstring, Change format of nested functions.
* ⚡ specialize edge-case temperature=inf in softmin_sentence_score
* ♻️ simplify temperature lookup
* ♻️ cleanup get_label_quality_scores
Remove unused args/variables. Update parameter list in docstring. Rename parameter of inner function. Function always returns a tuple.
* 🏷️ tag token_scores as optional
* ✅ test raised error
* 🩹 skip untestable elif statement
the elif statement only gets partial coverage because it can't evaluate to False due to the `assert sentence_score_method` at the start of the function
* 🏷️ restrict parameters for list types
* 🐛 only process characters in input token
Example: process_token("Cleanlab", [("C", "a"), ("a", "C")]) should return "aleCnlCb", not "CleCnlCb".
* ✅ use all sentences in test_get_sentence
* ✅ add test cases to test_filter_sentence
* ✅ extent test cases in test_mapping
* 📝 clean up docstrings
Restrict arg types based on docstrings. Fix punctuation and typos. Add examples to docstring.
* ✅ split tests for filter_sentence
* ✅ extend test_merge_probs
* ✅ test merge_probs with ignored/normalized columns in probs
* ✅ extend test cases for get_sentences
* ⚰️ remove unused pandas import
* 🏷️ pass strict mypy check
We ignore np.max as it is untyped.
No issues found in token_classification_utils.py by running
```
mypy --install-types --non-interactive --strict cleanlab/internal/token_classification_utils.py
```
* 👷 add strict type-checking in CI
* 💚 use strict type-check for single file
* ⏪ remove strict type check in CI
* refactor: 🏷️ use np.ndarray type instead of npt.NDArray
* ⏪ go back to generic np.ndarray type
* ♻️ always return tuple in filter_sentence
Remove unused argument+docstring. Simplify relevant unit tests.
* 🔥 resolve comments on typing
Remove ignore-comments. Remove duplicate tag in docstring. Remove unused imports.
* 🔥 remove duplicate tag in docstring
* calculate most likely class error from subset
* use verbose to control warning prints
* clip minimum to 1e-6 to prevent division by zero
* add docstring
* Changed all instances of np.array in docstring to np.ndarray
* np.array is NOT a proper class name it is just a function to
create np.ndarrays and therefore should not be parameter class
* Black formatting compliance
* findlabelissues doc clarifications + arg order
* clarify return_indices_ranked_by specifies return
* moved multi_label up higher. kept ranked_by at top
Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>
* validation_func docstring
* torch,tf compatibility+tests
* keras test
* skip tests if python < 3.7
* pytorch numpy int bug on windows
* make tensorflow test work on windows
* move tf env variable setting
* pytorch test increase epochs
* install cpu-tensorflow on windows CI
* torch test optimizer to adam
* fix bugs in shuffled TF dataset
* dummy unit test for TF on windows
* dummy code for TF windows testing
* deal with np.int bug on windows
* remove windows debugging code
* docstrings for new functionality
* address merge conflicts
* reformat after merge
* addressed comments
* 🏷️ annotate function args and return values
Starting with the validation module:
- I think (X, y) might need some custom Union type to handle both numpy arrays and pandas dataframe, etc.
- All of the "assert" functions return None.
Ref #307
* refactor: 🏷️ swap npt.NDArray -> np.ndarray
np.ndarray seems more consistent with the rest of the repo.
Maybe it's necessary to go back to npt.NDArray when disallowing generics?
See numpy docs: https://numpy.org/devdocs/reference/typing.html#numpy.typing.NDArray
* refactor: 🔥 remove unused import
* 🏷️ unconstrain X input types
* 🐛 handle label type issues
- Have to restrict the output of labels_to_array to pass mypy checks.
- Returning the values of pd.Series isn't type-stable.
* 🏷️ include np.generic in arg-type union
* test: ✅ test labels_to_array
* 🏷️ add type aliases for X and y
* 🚨 ignore type-checks for pandas indexing assertions
CI typechecker runs on Python 3.10 which gives this error:
'cleanlab/internal/validation.py:125: error: No overload variant of "__getitem__" of "_iLocIndexerSeries" matches argument type "List[int]"'
It should be fine to let mypy ignore these expressions as they don't return anything.
* 🥅 specify errors to ignore
"type: ignore" doesn't pass strict mypy type-checks unless the specific errors are provided
* 🏷️ annotate label series to array
* 🗑️ change labels arg -> y in CleanLearning.fit()
Set `label` as an optional keyword-only argument.
Anyone still using it in this method should get a deprecation warning.
Fixes#281
* 📝 add note for y/labels in docstring
* 🥅 make y an optional positional arg.
Should now resolve deprecated signatures.
* 📝 labels -> y in module docstring
* ⏪ revert "label -> y deprecation"
This reverts commit ab319a0cca2cec715a84eb5f628bbab7706c5f9c.
This reverts commit 1b739002d848e1f0acb6390a666f6e695e25fcaa.
This reverts commit 88bb6c3bcca298dab414c3cb20101783d78d35d1.
This reverts commit d988e3c3932107e779598d02d8f16d7e6671e9e7.
* ✨ add y alias for labels
Resolves#281
* Added runnable package versioning for tutorials
Added quickstart at top of tutorials
* Addressed PR comments
* Fixed quickstart to include cleanlearning
* Set correct requirements.txt
* Updated y to labels
* Fixed y-> labels labeling issues
* Updated quickstart message/removed extra dependencies
* remove keras from package-versions
* more concise
* remove pathlib as explicit requirement
* remove pathlib version
* rewording
* true label -> given label
* remove venv from kernelspec
Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
* cleanlearning tips
* addessed comments
* add a few general tips at the end
* minor edits
* format pred_probs as variable type
* bug fix to pass lint (missing comma)
* fix format for linter
Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>
* Updating tutorials hyperlink to 2.0.0 release
* revert back to v.2.0.0 links instead of stable
Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
* Added outlier detection tutorial into docs
* Switched to pytorch model/dataset implementation
* Fixed image normalization
* Fixed outlier detection for test set
* Added outlier thresholding into tutorial.
* Cleaned output
* Removed unused imports and renamed notebook
* Added outliers notebook to PR
* Added quickstart into tutorial
* Best quickstart header
* Cleaned cell output
* Changed to use subset of original data for speed
* Fixed randomness
* Cleared outputs
* fixed metadata tags
* Fixed metadata
* Cleaned kernel and verified output
* Improved unit test
Changed labels references to classes where apropriate
The previous link took you to a tutorials page with lots of options including tutorials that were for data-centric AI workflows, and computing cross validated probs, etc. if this is supposed to be the very first thing users click on to get started, we want it to take them straight to a place where they can get started in 5 minutes, not a list of things where they have to figure out what to click next.
Updated to fix this, which also reduced the length :)