Download PDF

European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, Date: 2020/09/14 - 2020/09/18, Location: Ghent, Belgium

Publication date: 2021-02-25
Volume: 12458 Pages: 121 - 136
ISSN: 978-3-030-67661-2
Publisher: Springer

ECML PKDD 2020: Machine Learning and Knowledge Discovery in Databases

Author:

Soenen, Jonas
Dumancic, Sebastijan ; Blockeel, Hendrik ; Van Craenendonck, Toon ; Hutter, F ; Kersting, K ; Lijffijt, J ; Valera, I

Keywords:

Science & Technology, Technology, Physical Sciences, Computer Science, Artificial Intelligence, Computer Science, Information Systems, Mathematics, Applied, Computer Science, Mathematics, Active learning, Clustering, Semi-supervised learning, CONSTRAINTS

Abstract:

Constraint-based clustering leverages user-provided constraints to produce a clustering that matches the user's expectation. In active constraint-based clustering, the algorithm selects the most informative constraints to query in order to produce good clusterings with as few constraints as possible. A major challenge in constraint-based clustering is handling noise: the majority of existing approaches assume that the provided constraints are correct, while that might not be the case. In this paper, we propose a method to identify and correct noisy constraints in active constraint-based clustering. Our approach reasons probabilistically about the correctness of the user’s answers and asks additional constraints to corroborate or correct the suspicious answers. We demonstrate the method’s effectiveness by incorporating it into COBRAS, a state-of-the-art method for active constraint-based clustering. Compared to COBRAS and other active-constraint-based clustering algorithms, the resulting system produces better clusterings in the presence of noise.