文件历史

提交图

210 次代码提交

作者 SHA1 备注 提交日期
Jonas Mueller c0b4e58471 multiannotator explanation improvements (#570)
more clarification when to use active learning vs fixed dataset analysis functions
2022-12-22 00:35:18 -08:00
Hui Wen d62fad78a6 Fix text tutorial, use GBM model in tabular tutorial (#565)
Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
2022-12-13 13:11:11 -08:00
Jonas Mueller 3a6f33c1ff typo: probabilites -> probabilities (#557) 2022-12-05 21:36:56 -08:00
Jonas Mueller 7dd06d14a0 remove unnecessary TF warning suppression statements in tutorials (#544) 2022-11-26 11:25:58 -08:00
Jonas Mueller 17a82c32fc Public multilabel quality scores method + softmin aggregation + more tests (#542)
Co-authored-by: Elías Snorrason <eliassno@gmail.com>
2022-11-23 18:34:43 -08:00
Jonas Mueller d2d481be46 dependencies between classes: multilabel tutorial 2022-11-17 19:14:55 -08:00
Jonas Mueller c86c51dd23 multilabel tutorial fixes: install colab, set seed, text improvements (#537)
Co-authored-by: Aditya Thyagarajan <aditya1593@icloud.com>
2022-11-17 07:09:25 -08:00
Aditya Thyagarajan ab2679724e Add multilabel tutorial to index.rst (#533) 2022-11-14 13:33:12 +00:00
Elías Snorrason 871f08693c Fix outdates links to example notebook for multilabel classification (#534) 2022-11-14 13:28:00 +00:00
Aditya Thyagarajan c32335c789 Tutorial for multi-label classification (#517)
clarification re empty list label, thanks to @elisno


Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
2022-11-11 17:13:56 -08:00
Jonas Mueller aceaa4f25e Improve tutorials intro language/formatting (#526) 2022-11-06 00:34:54 -07:00
Jonas Mueller 00d7c2157c studio pointers (#510) 2022-10-25 17:51:00 -07:00
Hui Wen 738acae810 Point to format label function for multiannotator (#506) 2022-10-25 10:52:21 -07:00
Jonas Mueller 00781ce984 update paper links (#503)
Co-authored-by: huiwengoh <45724323+huiwengoh@users.noreply.github.com>
2022-10-18 00:01:00 -07:00
Ulyana d27478853f Asthetic fixes for OOD tutorial (#493) 2022-10-03 10:26:14 -07:00
Jonas Mueller 6b8cb91908 Outlier tutorial: move uninteresting code to hidden cell (#492) 2022-09-30 00:17:48 -07:00
Jonas Mueller aedd0da886 mention crowdlab in multiannotator tutorial (#491) 2022-09-30 00:05:49 -07:00
Hui Wen bee40afd38 Doc edits: change header size in tutorials and better supress tf errors (#484) 2022-09-28 21:27:53 -07:00
Jonas Mueller 34acd0f19f tf-io tutorial version pinned, docstring improvements [include in 2.1 docs] (#466) 2022-09-22 14:37:44 -07:00
Ulyana 5da046da9e Added support for returning ranked issue idxs (#459)
* Added support for returning ranked issue idxs
- code
- tests
- docstring

Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
Co-authored-by: Curtis G. Northcutt <curtis.northcutt@gmail.com>
2022-09-16 20:20:52 -07:00
Anish Athalye 02ab55dff0 Add missing backticks and language annotation (#461) 2022-09-16 18:52:36 -07:00
Jonas Mueller 2eec20ea44 remove moderate dimension comment 2022-09-16 15:14:00 -07:00
Anish Athalye d8cc2e5056 Fix details disclosure elements in docs (#456)
This patch makes the following fixes:

- Make summary text appear on the same line as the ">" arrow marker;
  prior to this patch, it was being pushed down to the next line because
  there were nested block elements inside the <summary> element
- Style details element to visually delineate collapsible content with
  dashed lines above and below the content
- Add syntax highlighting for code blocks in details disclosure elements
  (some were missing syntax highlighting)
- Hide a cell that was meant to be hidden in the audio tutorial

The summary element looked ugly due to a block element being present
inside the <summary>. This was present in the rendered docs even though
it was not explicitly in the source. Even though our source (in a
Markdown cell in a Jupyter notebook) looked like:

    <summary>Example text</summary>

The rendered docs had:

    <summary><p>Example text</p></summary>

The easy fix was to leave this as-is and fix the styling with CSS,
making all elements inside the summary to be displayed as inline
elements.

Where was the extra <p> tag coming from? Different Markdown parsers
handle mixed Markdown/HTML differently. We are using nbsphinx, which
does something especially weird [1]: converting from Markdown ->
reStructuredText -> HTML.

For example, consider the following Markdown:

    <details><summary>Summary</summary>

    Body.

    </details>

nbsphinx was first converting it to this reST:

    .. raw:: html

       <details>

    .. raw:: html

       <summary>

    Summary

    .. raw:: html

       </summary>

    Body.

    .. raw:: html

       </details>

And then this reST was being converted to HTML. This is where the extra
<p> was coming from.

One workaround for this behavior is to switch to raw HTML cells in the
notebook, rather than Markdown cells, and write pure HTML there. This
has the disadvantage that it's inconvenient when the content inside the
<details> has extra markup: Markdown is very convenient for that (e.g.,
for styling text), and having the contents being parsed as Markdown also
allows us to easily use syntax highlighting inside the body.

Another workaround is to use raw reST cells [2] and write the proper
reST directly (avoiding the bad Markdown -> reST conversion). This
preserves the ability to use convenient formatting (though with reST
instead of Markdown) as well as syntax-highlighted code blocks. For
example:

    .. raw:: html

        <details><summary>Summary</summary>

    .. code-block:: python

        def meaning_of_life():
            return 42

    .. raw:: html

        </details>

The above snippet produces the expected HTML, with syntax-highlighted
code in the body, and without the extra <p> inside the <summary>.

Rather than either of these workarounds, the most convenient one for
developers is to fix the issue using CSS, so that the Markdown cells in
the tutorials can be left as-is and no special care is required when
writing details disclosure elements, which is what this patch does.

[1]: https://github.com/spatialaudio/nbsphinx/blob/fc79fd1ac5ab6df225b256ef4ca203c58c5ad92b/src/nbsphinx.py#L1309
[2]: https://nbsphinx.readthedocs.io/en/0.8.9/raw-cells.html#reST
2022-09-16 12:11:05 -07:00
Jonas Mueller a24dda5da8 clarify features for multiannotator tutorial 2022-09-16 11:22:02 -07:00
Jonas Mueller 7256cd3546 More improvements to token classification code and documentation (#452)
* improved token docs/code

* format paragraphs in docstrings

* fix typo

Co-authored-by: Elías Snorrason <eliassno@gmail.com>
2022-09-16 14:04:45 +00:00
Jonas Mueller 688c39c202 hide redundant display_example in audio tutorial 2022-09-16 04:09:15 -07:00
Jonas Mueller 5d900200e5 mention make_data in multiannotator tutorial 2022-09-16 03:46:03 -07:00
Elías Snorrason 5167d23b71 Update links to example notebooks (#444)
Related to cleanlab/examples#21
2022-09-15 11:10:33 -04:00
Hui Wen d74dda0ed9 fix code and md inconsistency (#442) 2022-09-14 14:45:09 -07:00
Jonas Mueller 84af558e85 hide more code in image tutorial, organize imports 2022-09-13 12:39:43 -07:00
Jonas Mueller 6a0e3bc461 Improvements to token_classification tutorial (#419)
* edits to token tutorial

* fix typos

* given_labels -> labels in quickstart block

* add IOB2-formatting clarification

Co-authored-by: Elías Snorrason <eliassno@gmail.com>
2022-09-13 09:34:33 -07:00
Jonas Mueller b9ab4ba57b suppress tensorflow warning logs in tutorials if not properly installed (#432)
* remove ineffective TF suppression lines

* suppress tensorflow logs in workflows tutorial

* formatting

* hidden cell comment

* move suppress code to hidden cell

* move suppression code to first hidden cell

* move suppression code to first cell

* supress tf logs in tabular tutorial

* optional code to hidden cells in audio tutorial

* suppress tf logs in multiannotator tutorial

* suppress tf logs in token tutorial

* move cross-val tutorial to appear last

* move cross-val tutorial to be last in main index

* move outlier before multiannotator in API list

* remove newline at end of cell
2022-09-13 00:47:21 -07:00
Jonas Mueller 15899f8f20 try to suppress tensorflow logs pt2 (#431)
* move env variable setting before tf import

* suppress tf logs in faq
2022-09-12 18:03:00 -07:00
Jonas Mueller a0276ca587 Text tutorial improvements (#429)
* more tf supression

* formatting

* move imports to top
2022-09-11 14:46:04 -07:00
Elías Snorrason 9a109182b1 Update argument given_labels -> labels in token classification module (#423)
* given_labels -> labels

only applies to the argument to `display_issues`

Fixes #418

* update kwargs in test for display_issues

* apply black formatter
2022-09-11 12:38:38 -07:00
Jonas Mueller f7102db902 Add links to examples notebooks where appropriate (#428)
* links to OOD examples notebook

* iterative notebook title in its link
2022-09-11 12:24:51 -07:00
Jonas Mueller c307e9714f update workflows for post 2.0 (#427)
* update workflows for post 2.0

* formatting

* link for cleanlearning
2022-09-11 00:27:57 -07:00
Jonas Mueller 78e8496363 Small tutorial notebook fixes/improvements (#420) 2022-09-09 17:36:27 +01:00
Hui Wen d8147caded Polish multiannotator docs (#422)
* docstring formatting

* add tiebreak info

* add example link to tutorial

* add missing backticks

* minor docstring text edits

* clarify "task"

Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
2022-09-09 09:15:54 -07:00
Elías Snorrason 44cd416a96 add tutorial notebook to docs.cleanlab.ai (#411) 2022-09-08 18:27:28 -07:00
Elías Snorrason ce90af35c6 Review tutorial for token classification (#410)
* review tutorial for token classification

- update labels variable name in markdown block
- add newline for rendering bullet points
- put pulldown content into code blocks
- hide code blocks and move filepaths variable

* resolve pr comments

* remove pulldown code snippet and todo

* fix notebook json - delete extra comma

Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
2022-09-08 18:27:11 -07:00
Ulyana 9dfa001068 Implementing get_ood_scores function (#338)
* Added base structure outline for get_ood_scores()
Added base test for get_ood_scores()

* Added warning for illogical param combo

* Addressed PR comments

* Added better unit tests
* TODO: test for correctly identifying OOD example

* Moved logic from get_ood_scores to _subtract_confident_thresholds

* TODO: Is this a cleaner way of doing this? Minimal code repeat but strange change to _subtract function
* Addressed type issue

* Switched logic for getting confident thresholds

* Fixing typecheck issues wiht labels parameter being None

* Simplified helper function. Testing type

* Fixed mypy static typing issue

* Mypy typecheck logic test

* removed uncessesary imports in util file

* typechecker debugging (add assert)

* Fixed type logic and removed confident_thresholds=None return

* Added extra arg in helper func to end of func

* Added zero-index checking for label param

* Added ood examples to outlier score notebook

* Added skeleton file structure for implementing outliers

* Make adjust_pred_probs=True by default not false

* Added base Outlier class functionality

TODO:
* test_outlier.py

* Added logic tests for function

* Updated get_ood_scores to always return confident_thresholds
* Even if none were calculated (then return is None)

* Added warning for fit that doesn't calculate confident_thresholds

* Moved get_outlier_scores and get_ood_scores to outlier.py

* Changed param dicts to dicts

* Added docstring to outlier.py

* Added proper return types

* Fixed mypy typing issues

* Switched outliers -> features; ood -> predictions naming conv

* Switched docstring to stem from fit and score functions

* Changed return of helper functions

* Fixed tutorials notebook to use OutOfDistribution class

* Moved imports to top of file

* Added option for different knn objects, addressed pr comments

* Switched params arg to init only

* Addressed PR comment for notebook, cleared notebook

* Fixed PR comments, wording.

* Testing relative links on build

* Changed ood_pred_probs scores to reflect 0 = most ood 1 = least

* Added MLP for detection outliers with pred_probs

* Fixed typos

* Added MLP Import

* Improved warning when fit call unnecessary

* Changed referenced to params dict in warnings/errors

* Added bagging+MLP classifier into notebook

* Reverted tutorial wording

* OOD tutorial improvements

* cleanup OOD documentation

* adopting->using

Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
2022-09-07 10:52:55 -07:00
Jonas Mueller f77f25f5a2 re-add hidden cell metadata 2022-09-06 15:40:26 -07:00
Jonas Mueller 92427d10b3 set data to None to avoid warning 2022-09-06 14:06:41 -07:00
Jonas Mueller bc8de13927 also demonstrate cleanlearning with pred_probs
Thanks @elisno for the suggestion!
2022-09-06 14:01:58 -07:00
Jonas Mueller 1375a27beb faq on predicted labels (#402) 2022-09-06 13:35:30 -07:00
Jonas Mueller 6ab83fc37c language improvements (#380) 2022-08-30 16:38:56 -07:00
Hui Wen 960c2b4ac1 CL functionality for multiannotator data (#333)
* setup multiannotator functionality

* add docstring

* add unittests

* edit get label quality with nan helper func

* docstring edits

* handle non overlapping annotators in get_annotator_agreement_with_annotators

* return NaN annotator quality scores for non-overlapping annotators

* add unittest

* add unittests for get_consensus_label tiebreaks

* np.array -> np.ndarray

* add new annotator quality method

* docstring edits

* address comments, docstring changes

* add method to compute improved consensus labels

* separate detailed_label_quality return

* elif statement typo

* changed get_majority_vote_label

* add 'best_quality' as a consensus method

* change lqs kwargs naming

* change lqs kwargs naming

* change unittest function names

* docstring edits

* change get_worst_class to return class with lowest agreement with consensus

* edit method to get annotator lqs, change quality_of_consensus -> consensus_quality_score

* remove unnecessary args

* satisfy mypy checks

* fix mypy issue

* add unittests

* reshuffled order of functions

* address comments

* docstring edit

* change return to always return dict

* add clipping to prevent negative pred_probs

* add tutorial

* add multiannotator to tutorial index

* add newline

* add tiebreaks for best_quality labels and worst_class

* change auto method name to crowdlab

* fix indexing for non numeric anno names

* add test for nonnumeric anno names

* update unittest

* address comments [partially complete]

* edit metadata

* missing /details tag

* bugfix to allow pd Int64 types

* update tutorial

* address comments
2022-08-30 11:29:06 -07:00
Eric Wang 1bad2f82a2 Adding functionality for cleanlab to find label errors in token classification datasets (#347)
* add token_classification functionality

* add typing

* fix typing

* add typing

* fixed typing and black

* fix typing
2022-08-30 09:17:07 -07:00
Hui Wen bf83c553bf make pred_probs naming consistent in tutorials (#345)
* rename cv_pred_probs to pred_probs

* revert python version
2022-08-17 19:52:25 -07:00