Clustering-Based Approach for Data Anonymization

微信服务号

微信订阅号

2025-5-13- 1

Home > Archive>Volume 21, Issue 4, 2010 >680-693

Clustering-Based Approach for Data Anonymization
DOI:
                        
                    
Author:
                        WANG Zhi-HuiWANG Zhi-Hui

Find this author on CNKI
Find this author on BaiDu
Search for this author on this site
XU JianXU Jian

Find this author on CNKI
Find this author on BaiDu
Search for this author on this site
WANG WeiWANG Wei

Find this author on CNKI
Find this author on BaiDu
Search for this author on this site
SHI Bai-LeSHI Bai-Le

Find this author on CNKI
Find this author on BaiDu
Search for this author on this site

                    
Affiliation:
Clc Number:
Fund Project:

Article

Figures

Metrics

Reference

Cited by

Materials

Comments

Abstract:

To prevent the disclosure of privacy, it requires preserving the anonymity of sensitive attributes in data sharing. The attribute values on quasi-identifiers often have to be generalized before data sharing to avoid linking attack, and thus to achieve the anonymity in data sharing. Data generalization increases the uncertainty of attribute values, and results in the loss of information to some extent. Traditional data generalization is often based on the predefined hierarchy, which causes over-generalization and too much unnecessary information loss. In this paper, the attributes in a quasi-identifier are classified into two categories, ordered attributes and unordered attributes. More flexible strategies for data generalization are proposed for them, respectively. At the same time, the loss of information is defined quantitatively based on the change of uncertainty of attribute values during data generalization. Furthermore, data anonymization is modeled by a clustering problem with special constraints. A clustering-based approach, called L-clustering, is presented for the l-diversity model. L-clustering can meet the requirement of preserving anonymity of sensitive attributes in data sharing, and reduce greatly the amount of information loss resulting from data generalization for implementing data anonymization.

Key words:data anonymization; quasi-identifier; linking attack; clustering; information loss

Get Citation

王智慧,许俭,汪卫,施伯乐.一种基于聚类的数据匿名方法.软件学报,2010,21(4):680-693

Copy

Article Metrics

Abstract:
PDF:
HTML:
Cited by:

History

Received:June 21,2007
Revised:October 08,2008
Adopted:
Online:
Published:

You are the firstVisitors
Copyright: Institute of Software, Chinese Academy of Sciences Beijing ICP No. 05046678-4
Address：4# South Fourth Street, Zhong Guan Cun, Beijing 100190,Postal Code：100190
Phone：010-62562563 Fax：010-62562533 Email：jos@iscas.ac.cn
Technical Support：Beijing Qinyun Technology Development Co., Ltd.

Beijing Public Network Security No. 11040202500063

微信服务号

微信订阅号

Get Citation

Share

微信扫一扫：分享

Article Metrics

History