期刊
WORLD WIDE WEB-INTERNET AND WEB INFORMATION SYSTEMS
卷 20, 期 6, 页码 1507-1525出版社
SPRINGER
DOI: 10.1007/s11280-017-0449-x
关键词
Streaming data; Class imbalance; Multi-window; Ensemble learning
资金
- ARC DP project [DP 130101327]
- 973 Program [2013CB329601, 2013CB329602, 2013CB329604]
- 863 Program [2012AA01A401, 2012AA01A402]
Imbalanced streaming data is commonly encountered in real-world data mining and machine learning applications, and has attracted much attention in recent years. Both imbalanced data and streaming data in practice are normally encountered together; however, little research work has been studied on the two types of data together. In this paper, we propose a multi-window based ensemble learning method for the classification of imbalanced streaming data. Three types of windows are defined to store the current batch of instances, the latest minority instances, and the ensemble classifier. The ensemble classifier consists of a set of latest sub-classifiers, and the instances employed to train each sub-classifier. All sub-classifiers are weighted prior to predicting the class labels of newly arriving instances, and new sub-classifiers are trained only when the precision is below a predefined threshold. Extensive experiments on synthetic datasets and real-world datasets demonstrate that the new approach can efficiently and effectively classify imbalanced streaming data, and generally outperforms existing approaches.
作者
我是这篇论文的作者
点击您的名字以认领此论文并将其添加到您的个人资料中。
推荐
暂无数据