4.5 Article

A simple and effective outlier detection algorithm for categorical data

Journal

Publisher

SPRINGER HEIDELBERG
DOI: 10.1007/s13042-013-0202-4

Keywords

Outlier detection; Categorical data; Weighted density; Information entropy

Funding

  1. National Natural Science Foundation of China [71031006]
  2. Foundation of Doctoral Program Research of Ministry of Education of China [20101401110002]
  3. Construction Project of the Science and Technology Basic Condition Platform of Shanxi Province [2012091002-0101]
  4. Shanxi Scholarship Council of China [2013-101]

Ask authors/readers for more resources

Outlier detection is an important data mining task that has attracted substantial attention within diverse research communities and the areas of application. By now, many techniques have been developed to detect outliers. However, most existing research focus on numerical data. And they can not directly apply to categorical data because of the difficulty of defining a meaningful similarity measure for categorical data. In this paper, a weighted density definition is given firstly, which takes account of the density and uncertainty of objects in every attributes simultaneously. Furthermore, a simple and effective outlier detection algorithm for categorical data based on the given weighted density is proposed. The corresponding time complexity of the algorithm is analyzed as well. Experimental results on real and synthetic data sets demonstrate the effectiveness and efficiency of our proposed algorithm.

Authors

I am an author on this paper
Click your name to claim this paper and add it to your profile.

Reviews

Primary Rating

4.5
Not enough ratings

Secondary Ratings

Novelty
-
Significance
-
Scientific rigor
-
Rate this paper

Recommended

No Data Available
No Data Available