文件历史

提交图

13 次代码提交

作者 SHA1 备注 提交日期
Ulyana 980c6ddc07 Extend package support Python 10-14 + relax dependencies (#1276) 2026-01-07 14:11:40 -08:00
Jonas Mueller 0167accd0a license (#1263)
change license to Apache 2.0

---------

Co-authored-by: Anish Athalye <me@anishathalye.com>
2025-12-15 16:56:30 -08:00
Elías Snorrason ffdbe77dc6 Enhance knn-based outlier detection in Datalab (#1163) 2024-06-25 15:53:56 +00:00
Elías Snorrason 2a68fd7205 Improve knn graph handling and outlier detection in issue managers (#1155) 2024-06-24 23:49:33 +00:00
Elías Snorrason 71ba4b3209 Add a knn module (#1117)
Co-authored-by: Jonas Mueller <1390638+jwmueller@users.noreply.github.com>
2024-05-14 13:59:15 -04:00
Elías Snorrason 2cd706976f Use a safer value of median_nn_distance for division in duplicate.py (#1116) 2024-05-04 04:02:53 +00:00
Elías Snorrason b13d27e9b9 Fix numerical instability with Euclidean distance metric (#1113) 2024-05-02 13:24:58 +00:00
Elías Snorrason 2f2bc1fe18 Refine Scoring and Enhance Stability for Datasets with Identical Examples (#1056)
Enhance numerical stability and scoring accuracy for datasets with identical features

- Adjust minimum median values to 100x machine epsilon to prevent unstable operations.
- Implement specific scoring logic for outlier and near-duplicate issues in datasets with identical features, ensuring correct scores are assigned in such edge cases.
- Introduce unit tests to validate the expected behavior for datasets comprising entirely identical examples. Check an additional case with an extra unique example.

Related to issue #1055.
2024-03-19 12:41:49 +00:00
Elías Snorrason 23af72a2bd Near duplicate score transformation (#943) 2024-01-05 01:04:16 +00:00
Ryan Singman bc10585352 Patch: epsilon for near duplicate (#919)
Co-authored-by: Elías Snorrason <eliassno@gmail.com>
2023-12-21 13:03:51 +00:00
Elías Snorrason e9ee35f075 Test properties of near duplicate sets (#895)
* create and test strategy for generating knn_graph as test inputs

* test strategy for generating knn_graphs

- Each row has the distances sorted in ascending order in the csr-format.
- The indices of the neighbors are unique within a column and don't have the query point as a neighbor.
- If points a and b are mutual neighbors, they have the same distance between them.

* add property based tests for near-duplicate sets

* flag near-duplicate issues based on items in near duplicate sets.
2023-11-28 15:06:48 +00:00
Tata Ganesh 62980c1334 Catch NotFitted exception for knn (#825) 2023-08-28 14:08:02 +00:00
Elías Snorrason b0fb8e5f77 Move Datalab internals to dedicated internal module (#783) 2023-07-28 16:18:51 +00:00