☆ 4.7 Article

Entity reconciliation in big data sources: A systematic mapping study

EXPERT SYSTEMS WITH APPLICATIONS (2017)

Journal

EXPERT SYSTEMS WITH APPLICATIONS

Volume 80, Issue -, Pages 14-27

Publisher

PERGAMON-ELSEVIER SCIENCE LTD

DOI: 10.1016/j.eswa.2017.03.010

Keywords

Systematic mapping study; Entity reconciliation; Heterogeneous databases; Big data

Funding

MeGUS project [TIN2013-46928-C3-3-R]
Pololas project [TIN2016-76956-C3-2-R]
SoftPLM Network of the Spanish the Ministry of Economy and Competitiveness [TIN2015-71938-REDT]
Fujitsu Laboratories of Europe (FLE)

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Abstract

The entity reconciliation (ER) problem aroused much interest as a research topic in today's Big Data era, full of big and open heterogeneous data sources. This problem poses when relevant information on a topic needs to be obtained using methods based on: (i) identifying records that represent the same real world entity, and (ii) identifying those records that are similar but do not correspond to the same real-world entity. ER is an operational intelligence process, whereby organizations can unify different and heterogeneous data sources in order to relate possible matches of non-obvious entities. Besides, the complexity that the heterogeneity of data sources involves, the large number of records and differences among languages, for instance, must be added. This paper describes a Systematic Mapping Study (SMS) of journal articles, conferences and workshops published from 2010 to 2017 to solve the problem described before, first trying to understand the state-of-the-art, and then identifying any gaps in current research. Eleven digital libraries were analyzed following a systematic, semiautomatic and rigorous process that has resulted in 61 primary studies. They represent a great variety of intelligent proposals that aim to solve ER. The conclusion obtained is that most of the research is based on the operational phase as opposed to the design phase, and most studies have been tested on real-world data sources, where a lot of them are heterogeneous, but just a few apply to industry. There is a clear trend in research techniques based on clustering/blocking and graphs, although the level of automation of the proposals is hardly ever mentioned in the research work. (C) 2017 Elsevier Ltd. All rights reserved.

Entity reconciliation in big data sources: A systematic mapping study

Journal

EXPERT SYSTEMS WITH APPLICATIONS

Publisher

PERGAMON-ELSEVIER SCIENCE LTD

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Entity reconciliation in big data sources: A systematic mapping study

Journal

EXPERT SYSTEMS WITH APPLICATIONS

Publisher

PERGAMON-ELSEVIER SCIENCE LTD

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper