Our package doesn't have type annotations everywhere, so we can't use
mypy in strict mode just yet. Still, adding type checking in CI is
valuable, so we don't have unchecked annotations in our code.
This patch includes basic fixes to make type checking pass, including
switching the incorrect `np.array` type annotation for `np.ndarray` and
adding some assertions for flow-sensitive typing.
* cleanlearning w dfs
* work for sparse matrix as well
* simplify logic of labels_to_array and extend types
* address pr feedback
* add unit test
* rare label dataframe
* modularize subsetting code
* series rarelabel test
* replace cal.com with slack/email
* Add general method to find num_classes from labels
* compute num_classes with pred_probs.shape[1]
* fix broken commits, address 2nd round of comments
Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>
Add clarification of the labels format requirements for all major API functions.
* Fix broken link
* Clarify reqs for labels format. rank.py does not yet support multi_label
* add double ticks to code in docstrings
* clarify docstring
* further clarify multi_label vs single label labels reqs
* also add new docstring to confident joint
* Add KNN distance OOD scoring function and unit tests
* Update KNN distance OOD scoring function
* Change query_features to features in unit tests for KNN distance OOD scoring function
* Update KNN distance OOD scoring function
* Update tests for KNN distance OOD scoring function to use auto for algo
* Allow k=None for KNN Distance OOD score
* added class to use cleanlab with tensorflow and huggingface models
* - Documentation refactoring
- predict and predict_proba functions work on new test data too
* Add log_loss_search weighting method for ensemble label quality scoring function
* Update log_loss_search weighting method
* Add test for log_loss_search method
* Add parameter to Ensemble label quality scoring function for t values in log_loss_search method
* Update ensemble label quality scoring function docstring
* Update ensemble label quality scoring function comments
* Update docstring in ensemble label quality scoring function
* Update docstring in ensemble label quality scoring function
* Update docstring in ensemble label quality scoring function
* Modify verbose printout for log_loss_search
* Add clipping of pred_prob when calculating weights for log_loss_search
* Add clipping of pred_prob and renormalization when calculating weights for log_loss_search
* Add comments for log_loss_search weighting scheme
* Allow user to pass custom_weights to ensemble scoring method
* Add tests for ensemble scoring with custom_weights
* Add check to make sure length of custom_weights matches len(pred_probs_list)
* Update tests for usage of custom_weights in ensemble scoring