* cleanlearning w dfs
* work for sparse matrix as well
* simplify logic of labels_to_array and extend types
* address pr feedback
* add unit test
* rare label dataframe
* modularize subsetting code
* series rarelabel test
* replace cal.com with slack/email
* Add general method to find num_classes from labels
* compute num_classes with pred_probs.shape[1]
* fix broken commits, address 2nd round of comments
Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>
Add clarification of the labels format requirements for all major API functions.
* Fix broken link
* Clarify reqs for labels format. rank.py does not yet support multi_label
* add double ticks to code in docstrings
* clarify docstring
* further clarify multi_label vs single label labels reqs
* also add new docstring to confident joint
* Add KNN distance OOD scoring function and unit tests
* Update KNN distance OOD scoring function
* Change query_features to features in unit tests for KNN distance OOD scoring function
* Update KNN distance OOD scoring function
* Update tests for KNN distance OOD scoring function to use auto for algo
* Allow k=None for KNN Distance OOD score
* added class to use cleanlab with tensorflow and huggingface models
* - Documentation refactoring
- predict and predict_proba functions work on new test data too
* Add log_loss_search weighting method for ensemble label quality scoring function
* Update log_loss_search weighting method
* Add test for log_loss_search method
* Add parameter to Ensemble label quality scoring function for t values in log_loss_search method
* Update ensemble label quality scoring function docstring
* Update ensemble label quality scoring function comments
* Update docstring in ensemble label quality scoring function
* Update docstring in ensemble label quality scoring function
* Update docstring in ensemble label quality scoring function
* Modify verbose printout for log_loss_search
* Add clipping of pred_prob when calculating weights for log_loss_search
* Add clipping of pred_prob and renormalization when calculating weights for log_loss_search
* Add comments for log_loss_search weighting scheme
* Allow user to pass custom_weights to ensemble scoring method
* Add tests for ensemble scoring with custom_weights
* Add check to make sure length of custom_weights matches len(pred_probs_list)
* Update tests for usage of custom_weights in ensemble scoring
* df return type, need tests still
* Add pandas as a dependency
We already decided that pandas will be a dependency of cleanlab (also
used in the dataset module, see
https://github.com/cleanlab/cleanlab/pull/182).
* Tweak documentation
* addressed comments
* remove lazy import
* address 2nd round comments
* unit tests
* improve codecov
* Fix typo
* methods to save more space
* nocover statements for prints
* extra nocover
* nocover warnings
* test docstring formatting
* test docstring formatting2
* test docstring formatting2
* move compress to helper, find-label docs params
* readded stuff lost in merge conflict
* addressed remaining PR review comments
* docs formatting
* docs formatting2
* docs formatting3
* docs formatting4
* docs formatting5
* docs formatting5
* docs formatting6
* docs formatting7
* docs formatting8
* docs formatting9
* docs formatting19
* docs formatting20
* docs formatting20
* docs formatting21
* code formatting
* fix a bug where confident joint isnt computed
The confident joint wasn't getting computed if noise_matrix was passed in and pred_probs was not passed in. But that's bad because it stops workflows like:
```python
cl = CleanLearning()
cl.fit(data, labels, noise_matrix=noise_matrix)
cleanlab.dataset.health_summary(labels, confident_joint=cl.confident_joint)
```
* fixed bug from last commit. code in wrong place.
* print overwrite bugfix
Co-authored-by: Anish Athalye <me@anishathalye.com>
Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>