* Changed all instances of np.array in docstring to np.ndarray
* np.array is NOT a proper class name it is just a function to
create np.ndarrays and therefore should not be parameter class
* Black formatting compliance
* findlabelissues doc clarifications + arg order
* clarify return_indices_ranked_by specifies return
* moved multi_label up higher. kept ranked_by at top
Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>
Add clarification of the labels format requirements for all major API functions.
* Fix broken link
* Clarify reqs for labels format. rank.py does not yet support multi_label
* add double ticks to code in docstrings
* clarify docstring
* further clarify multi_label vs single label labels reqs
* also add new docstring to confident joint
* Makes some methods private that are not intended to be user-facing.
* Adds experimental module with fasttext.py and coteaching.py
* Adds header descriptions to code files which will render in docs
* Many miscellaneous fixes
* added LearningWithNoisyLabels.find_label_issues instance method
* LNL.find_label_issues no longer memoizes
* verbose unit test coverage
* Fixed all issues in PR. Fixed confident joint usage in LNL.find_label_issues. Fixed other minor bugs. Added tests
Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>
Co-authored-by: Anish Athalye <me@anishathalye.com>
* improved language in docs Quickstart
* LearningWithNoisyLabels refactor to improve API flexibility and UX
* fix LearningWithNoisy rare label handling
* additional arg-checking unit tests
* utilities -> internals
* docbuilding instructions improvement
* fixup formatting of contributing.md
* change contributor guidelines language to be optional
* line formatting
* Add label quality scoring function for users to choose the scoring method.
Add confidence_weighted_entropy as another label quality score that is available for scoring and ranking.
Add option in scoring function to adjust predicted probabilities by subtracting the class thresholds.
Refactor order_label_issues to use the new label quality scoring function and accept keyword args.
* Cleanup docstring in label quality scoring functions
* Change rank_by_kwargs default to empty dict. Dict is used as keyword args for label quality scoring function.
* Cleanup docstrings. Add **rank_by_kwargs to signature of order_label_issues function.
* Add exception handler for invalid rank_by methods
* Add **rank_by_kwargs to allow keyword args in find_label_issues()
* Add test for confidence_weighted_entropy rank scoring function
* Add test for scoring function that accepts scoring method
* Cleanup comments
* Update test for scoring functions
* Update test for scoring functions
* Modify order_label_issues() function to run score_label_quality with (labels, pred_probs) and then filter with label_issues_mask. This is more robust to allow us to adjust the pred_probs (e.g. subtract confident class thresholds)
* Update test for label quality scoring
* Update test for label quality scoring
* Update exception handler for label scoring function
* Add test for subtracting confident class thresholds
* Update test for subtracting confident class thresholds
* Update test for subtracting confident class thresholds
* Add ensemble label quality scoring function
* Update test for ensemble label quality scoring function
* Update test for ensemble label quality scoring function
* Update docstrings for functions that accept adj_pred_probs with description of the adjustment to pred_probs
* Update find_label_issues to pass rank_by_kwargs as a dict
* Cleanup comments
* Refactor tests for scoring functions to avoid redundancies
* Refactor tests for scoring functions to avoid redundancies. Use parameterize instead of for-loop
* Parameterize scoring method to test ensemble scoring function
* Print scoring method when scoring function test fails
* Cleanup docstrings to avoid redundancies
* Move class confident threshold functions to new label quality utils module. Add weight_ensemble_members_by parameter to label quality ensemble scoring function allows users to choose weighting scheme (uniform, accuracy). Enhance tests for CI.
* Add test for CI. Test bad arg for weight_ensemble_members_by parameter.
* Add explicit error checking of pred_probs_list arg.
* Cleanup docstring and print statements
* Cleanup docstrings. Move get_entropy to utils. Refactor score_label_quality_ensemble to print accuracy weights
* Change default label quality scoring method to normalized_margin. Add to docstring explaining when to use normalized_margin vs self_confidence.
* Cleanup docstring
* Cleanup docstrings. Rename get_entropy() to get_normalized_entropy()
* Rename score_label_quality() to get_label_quality_scores(). Rename score_label_quality_ensemble() to get_label_quality_ensemble_scores().
* Remove extra indent in docstring
* Rename adj_pred_probs to adjust_pred_probs. Move get_confident_thresholds() to count.py module.
* Add comment to explain why we add a dimension for numpy broadcasting.
* Change renormalization logic when adjusting pred_probs. Raise ValueError when adjust_pred_probs is used with unsupported scoring method. Add test for ValueError.
* Update tests to account for unsupported method when running adjust_pred_probs.
* Enhance docstring tests
* Refactor modules pruning to filter and latent_estimation to count
* Remove polyplex (research) algorithms from cleanlab
* Create new module rank and move scoring functions to rank.
* Rename test to match new module names
* Fixed error in normalized margin. added ranking for arbitrary psx and labels.
* Remove unused tests and methods. add multi-label support for baseline.
* Move baseline methods to filter and delete baseline module.
* change filter.get_noise_indices to filter.find_label_issues
* Rename baseline methods. fill out docstrings.
* Only require 1 example to be left in each class after removing errors. (instead of 5)
* Remove K as a parameter to count.compute_confident_joint
* Add C_argmax and C_ij methods from CL paper to find_label_issues
* Add warnings for new prune methods and frac_noise. Fix tests.
* Add baseline tests to test_rank_filter and delete baseline test
* Remove inverse_noise_matrix parameter in classification call to find_label_issues
* add todo to update docstring with new ranking functions
* 100% tests pass. add multi-label support for prune_method
* Major NOT-backwards-compatible name changes to most components
* More Major NOT-backwards-compatible name changes
* fixed s -> label mistakes
* Several nomenclature updates from PR feedback. models renamed to example models.
* Remove python2 support across all modules.
* major api changes. psx -> pred_probs. prob_given_label -> self_confidence. testing added.
* enable python version 3.9 for pytorch model.
* ran spellcheck
* ran grammar check
* Update count.py
* Update filter.py
* Update setup.py and ci.yml to no longer support Python 2 and py3.4/5
* Increase test coverage and documentation of rank module methods.
* create utils submodule and move util and latent_algebra
* Rename y everywhere to true_labels, and p(true_label=..)
* Enforce positional arguments in methods. Fully remove py2 support.