☆ 4.7 Article

Handling data irregularities in classification: Foundations, trends, and future challenges

PATTERN RECOGNITION (2018)

期刊

PATTERN RECOGNITION

卷 81, 期 -, 页码 674-693

出版社

ELSEVIER SCI LTD

DOI: 10.1016/j.patcog.2018.03.008

关键词

Data irregularities; Class imbalance; Small disjuncts; Class-distribution skew; Missing features; Absent features

类别

Computer Science, Artificial Intelligence Engineering, Electrical & Electronic

资金

Indian National Academy of Engineering (INAE)

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

摘要

Most of the traditional pattern classifiers assume their input data to be well-behaved in terms of similar underlying class distributions, balanced size of classes, the presence of a full set of observed features in all data instances, etc. Practical datasets, however, show up with various forms of irregularities that are, very often, sufficient to confuse a classifier, thus degrading its ability to learn from the data. In this article, we provide a bird's eye view of such data irregularities, beginning with a taxonomy and characterization of various distribution-based and feature-based irregularities. Subsequently, we discuss the notable and recent approaches that have been taken to make the existing stand-alone as well as ensemble classifiers robust against such irregularities. We also discuss the interrelation and co-occurrences of the data irregularities including class imbalance, small disjuncts, class skew, missing features, and absent (non-existing or undefined) features. Finally, we uncover a number of interesting future research avenues that are equally contextual with respect to the regular as well as deep machine learning paradigms. (C) 2018 Elsevier Ltd. All rights reserved.

Handling data irregularities in classification: Foundations, trends, and future challenges

期刊

PATTERN RECOGNITION

出版社

ELSEVIER SCI LTD

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

Handling data irregularities in classification: Foundations, trends, and future challenges

期刊

PATTERN RECOGNITION

出版社

ELSEVIER SCI LTD

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文