EEMCS

Home > Publications
Home University of Twente
Education
Research
Prospective Students
Jobs
Publications
Intranet (internal)
 
 Nederlands
 Contact
 Search
 Organisation

EEMCS EPrints Service


15237 Probabilistic Data Integration
Home Policy Brochure Browse Search User Area Contact Help

van Keulen, M. (2009) Probabilistic Data Integration. (Invited) In: 08421 Abstracts Collection - Uncertainty Management in Information Systems, 12 - 17 Oct 2008, Dagstuhl, Germany. pp. 8-8. Dagstuhl Seminar Proceedings (08421). Schloss Dagstuhl - Leibniz-Zentrum fuer Informatik. ISSN 1862-4405

Full text available as:

PDF

202 Kb

Official URL: http://drops.dagstuhl.de/opus/volltexte/2009/1942

Exported to Metis

Abstract

In data integration efforts such as in portal development, much development time is devoted to entity resolution. Often advanced similarity measurement techniques are used to remove semantic duplicates or solve other semantic conflicts. It proofs impossible, however, to automatically get rid of all semantic problems. An often-used rule of thumb states that about 90% of the development effort is devoted to semi-automatically resolving the remaining 10% hard cases. In an attempt to significantly decrease human effort at data integration time, we have proposed an approach that strives for a 'good enough' initial integration which stores any remaining semantic uncertainty and conflicts in a probabilistic XML database. The remaining cases are to be resolved during use with user feedback.
We conducted extensive experiments on the effects and sensitivity of rule denition, threshold tuning, and user feedback on the integration quality. We claim that our approach indeed reduces development effort - and not merely shifts the effort - by showing that setting rough safe thresholds and defining only a few rules suffices to produce a 'good enough' integration that can be meaningfully used, and that user feedback is effective in gradually improving the integration quality.

Item Type:Conference or Workshop Paper (Abstract, Invited/Keynote Talk)
Research Group:EWI-DB: Databases
Research Program:CTIT-NICE: Natural Interaction in Computer-mediated Environments
Research Project:MultimediaN/N3: Multimedia databases
Uncontrolled Keywords:Uncertainty management, Data integration, entity resolution, probabilistic databases, data quality
ID Code:15237
Status:Published
Deposited On:31 March 2009
Refereed:No
International:Yes
More Information:statisticsmetis

Export this item as:

To correct this item please ask your editor

Repository Staff Only: edit this item