Article ID: | iaor20123444 |
Volume: | 195 |
Issue: | 1 |
Start Page Number: | 97 |
End Page Number: | 110 |
Publication Date: | May 2012 |
Journal: | Annals of Operations Research |
Authors: | Abril Daniel, Navarro-Arribas Guillermo, Torra Vicen |
Keywords: | security |
Record linkage is used in data privacy to evaluate the disclosure risk of protected data. It models potential attacks, where an intruder attempts to link records from the protected data to the original data. In this paper we introduce a novel distance based record linkage, which uses the Choquet integral to compute the distance between records. We use a fuzzy measure to weight each subset of variables from each record. This allows us to improve standard record linkage and provide insightful information about the re‐identification risk of each variable and their interaction. To do that, we use a supervised learning approach which determines the optimal fuzzy measure for the linkage.