Towards Reliable Rare Category Analysis on Graphs via Individual Calibration

Wu, Longfeng; Lei, Bowen; Xu, Dongkuan; Zhou, Dawei

Towards Reliable Rare Category Analysis on Graphs via Individual Calibration

dc.contributor.author	Wu, Longfeng	en
dc.contributor.author	Lei, Bowen	en
dc.contributor.author	Xu, Dongkuan	en
dc.contributor.author	Zhou, Dawei	en
dc.date.accessioned	2023-09-05T13:39:14Z	en
dc.date.available	2023-09-05T13:39:14Z	en
dc.date.issued	2023-08-06	en
dc.date.updated	2023-09-01T07:49:34Z	en
dc.description.abstract	Rare categories abound in a number of real-world networks and play a pivotal role in a variety of high-stakes applications, including financial fraud detection, network intrusion detection, and rare disease diagnosis. Rare category analysis (RCA) refers to the task of detecting, characterizing, and comprehending the behaviors of minority classes in a highly-imbalanced data distribution. While the vast majority of existing work on RCA has focused on improving the prediction performance, a few fundamental research questions heretofore have received little attention and are less explored: How confident or uncertain is a prediction model in rare category analysis? How can we quantify the uncertainty in the learning process and enable reliable rare category analysis? To answer these questions, we start by investigating miscalibration in existing RCA methods. Empirical results reveal that stateof- the-art RCA methods are mainly over-confident in predicting minority classes and under-confident in predicting majority classes. Motivated by the observation, we propose a novel individual calibration framework, named CaliRare, for alleviating the unique challenges of RCA, thus enabling reliable rare category analysis. In particular, to quantify the uncertainties in RCA, we develop a node-level uncertainty quantification algorithm to model the overlapping support regions with high uncertainty; to handle the rarity of minority classes in miscalibration calculation, we generalize the distribution-based calibration metric to the instance level and propose the first individual calibration measurement on graphs named Expected Individual Calibration Error (EICE). We perform extensive experimental evaluations on real-world datasets, including rare category characterization and model calibration tasks, which demonstrate the significance of our proposed framework.	en
dc.description.version	Published version	en
dc.format.mimetype	application/pdf	en
dc.identifier.doi	https://doi.org/10.1145/3580305.3599525	en
dc.identifier.uri	http://hdl.handle.net/10919/116205	en
dc.language.iso	en	en
dc.publisher	ACM	en
dc.relation.ispartof	KDD '23: Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining	en
dc.rights	In Copyright	en
dc.rights.holder	The author(s)	en
dc.rights.uri	http://rightsstatements.org/vocab/InC/1.0/	en
dc.title	Towards Reliable Rare Category Analysis on Graphs via Individual Calibration	en
dc.type	Article - Refereed	en
dc.type.dcmitype	Text	en

Files

Original bundle

Now showing 1 - 1 of 1

Name:: 3580305.3599525.pdf
Size:: 2.13 MB
Format:: Adobe Portable Document Format
Description:: Published version

Download

License bundle

Now showing 1 - 1 of 1

Name:: license.txt
Size:: 0 B
Format:: Item-specific license agreed upon to submission
Description:

Download

Collections

Journal Articles, Association for Computing Machinery (ACM)
Scholarly Works, Computer Science