Research Data Pseudonymisation and Anonymisation: A Practical Guide
Abstract
The 'Research Data Pseudonymisation and Anonymisation: A Practical Guide' has been created by the Data Stewards from the Behavioural, Management, and Social Science (BMS) faculty and the Geo-Information Science and Earth Observation (ITC) faculty of the University of Twente. The guide provides practical guidance on two methods of data de-identification: data pseudonymisation and anonymisation. Data de-identification refers to the process of removing or minimising information that can be used to track back to a research participant within a dataset. This guide includes hands-on examples of various techniques for pseudonymising and anonymising different types of research data.
Full text
RESEARCH DATA PSEUDONYMISATION AND ANONYMISATION A PRACTICAL GUIDE
1 COLOPHON ORGANISATION Library, ICT Services & Archive (LISA) TITLE Research Data Pseudonymisation and Anonymisation: A Practical Guide VERSION 1.0 AUTHOR(S) (in alphabetical order) Minsi Li, Deniece S. Nazareth, Alice Nikuze DATE 31-10-2025 COPYRIGHT CC-BY
2 Disclaimer: The authors cannot be held responsible for any unintended negative consequences resulting from the incompleteness of this guide, or from the misuse or misinterpretation of its content or discrepancies with any other institutions’ guidance. PREFACE This guide provides practical guidance on two methods of data de-identification: data pseudonymisation and anonymisation. Data de-identification refers to the process of removing or minimising information that can be used to track back to a research participant within a dataset. This guide includes hands-on examples of various techniques for pseudonymising and anonymising different types of research data. The guide is structured as follows: Chapter 1 defines personal data and explains different types of identifiers. Chapter 2 introduces fundamental principles and concepts of data de-identification. It concludes with practical considerations and best practices. Chapter 3 delves into the techniques of pseudonymisation and anonymisation for textual, numerical, audio-visual, and geospatial research data. Additionally, practical examples are included. Chapters 4, 5 and 6 present the resources, glossary, and attribution. Although extensive techniques and examples are presented, please note that this guide does not cover every issue related to pseudonymisation or anonymisation and does not exhaust all types of research data. Furthermore, the de-identification of research data is often complex and requires a case-by-case approach. Therefore, please always consult your data steward or privacy contact person for more information.
3 TABLE OF CONTENTS Preface ..................................................................................................................... 2 1. Introduction ....................................................................................................... 4 1.1 Personal data in research ............................................................................. 4 1.1.1 Categories of identifiers ............................................................................ 4 1.1.2 Special categories of personal data ........................................................... 5 2. Principles, concepts and practical considerations ............................................ 6 2.1 Data minimisation ........................................................................................ 6 2.2 Pseudonymisation versus Anonymisation ..................................................... 7 2.2.1 Pseudonymisation .................................................................................... 7 2.2.2 Anonymisation .......................................................................................... 8 2.3 Data de-identification process ...................................................................... 9 2.4 Practical considerations ............................................................................. 11 3. Pseudonymisation and anonymisation techniques .......................................... 13 3.1 Textual data ............................................................................................... 13 3.2 Numerical data .......................................................................................... 16 3.3 Audio-visual data ....................................................................................... 21 3.4 Geospatial data .......................................................................................... 23 4. Resources ........................................................................................................ 25 5. Glossary ........................................................................................................... 26 6. Attribution ........................................................................................................ 28
4 1. INTRODUCTION In today's data-driven world, privacy concerns have become paramount. With the increasing collection of personal data in research, researchers bear the responsibility to protect the privacy of the data subjects. More specifically, the General Data Protection Regulation (GDPR) requires researchers affiliated with any institutions within the European Economic Area (EEA) to protect the privacy of research participants when their personal data is processed in research. In this context, researchers should employ all necessary data protection measures, including de-identification techniques, to minimise the risk of tracking back to a data subject. 1.1 PERSONAL DATA IN RESEARCH The term personal data refers to any information that enables people who have access to the data (including you as a researcher) to identify a living person. According to the GDPR, a living person can be identified, directly or indirectly, by means of an identifier such as a name, a citizen identification number, location data, or by other details specific to the person’s physical, physiological, genetic, mental, economic, cultural, or social identity. 1.1.1 Categories of identifiers There are three main categories of identifiers: • Direct identifiers: any piece of information unique to an individual and sufficient on its own to identify that individual, such as a name, a phone number, a picture, security numbers, audio/sound, video, photographs, etc. • Indirect identifiers: any information that can reveal someone's identity when combined with other available information. Examples include an address, age, gender, occupation, location, health-related records, ethnic group, etc. • Strong indirect identifiers: it is worth noting that some identifiers are particularly considered strong indirect identifiers because they can significantly increase the risk of easily identifying an individual when combined with other information. Examples of strong identifiers include a rare event, a rare disease, an unusual job title, a date of birth, an IP address, etc. See Figure 1 for examples of identifiers in the three main categories (next page).
5 Figure 1: Examples of identifiers typically processed in research There are other types of identifiers that can significantly increase the risks of reidentification. Notable examples are: • Biometric identifiers: which include unique and often permanent traits, such as physiological and behavioural characteristics, including facial features, fingerprints, DNA, voice, and gait. Biometric identifiers also include less distinctive traits that may change over time, such as visible marks (e.g., tattoos or scars). • Contextual elements: like speaking or writing styles, clothing style and hair style, and specific social or environmental settings can, in some cases, also increase the risks of identification. 1.1.2 Special categories of personal data There are special categories of personal data, also known as sensitive personal data. Sensitive personal data includes any information about a person’s racial or ethnic origin, political opinions, religious or philosophical beliefs, trade union membership, genetic data, biometric data, health, sex life, or sexual orientation. This personal data can put individuals or particular groups in vulnerable positions, and their misuse can potentially lead to physical or mental harm, financial loss, political or religious persecution, disclosure of sexual orientation, discrimination, or embarrassment. Therefore, handling such data requires extra caution.
6 2. PRINCIPLES, CONCEPTS AND PRACTICAL CONSIDERATIONS 2.1 DATA MINIMISATION Data minimisation is a fundamental principle of data privacy and protection. This principle requires limiting personal data collection and retention to only what is relevant, necessary, and adequate to accomplish a given research purpose. Collecting only the minimum amount of personal data necessary for your research reduces the risks associated with excessive data collection. More personal data implies higher risks of data breaches and their misuse, thereby increasing potential harm for the research participants. Personal data minimisation in research can be achieved by: • Carefully planning what personal data is necessary for your research purpose. Plan what specific personal data is truly needed early in the research design phase. For example, avoid collecting participants' ages if your study does not examine agerelated effects. • Collecting aggregated information when detailed information is not required. If broader categories (e.g., age ranges, geographic area) suffice for your analysis, there is no need to collect exact values. • Limiting open-ended questions in surveys. Open-ended questions can give research participants room to share unsolicited personal data about themselves or others. Reducing the number of open-ended questions can minimise the risks of collecting unnecessary personal data. • Instructing research participants sufficiently to avoid unnecessary personal data. Inform participants ahead of the data collection to refrain from sharing unnecessary personal information, such as names or addresses of others. Adding clear instructions and reminding the participants not to mention personal data (e.g., other people's names or addresses) can prevent unintentional personal data. • Adjusting the settings of the data collection tool to minimise personal data collection. Review and adjust the settings of data collection tools like survey platforms to prevent the default collection of unnecessary personal data, such as location or IP addresses. When data collection is finalised, please migrate personal data from the data collection tools to safe data storage as soon as possible.
7 2.2 PSEUDONYMISATION VERSUS ANONYMISATION When personal data are necessary and have been collected for research, research participants’ identities should be protected. Two techniques, pseudonymisation and anonymisation, are used to ensure that research participants remain anonymous and untraceable during and after research. 2.2.1 Pseudonymisation Pseudonymisation is a de-identification process after which research data can no longer be linked to a specific individual without additional information. This process also referred to as coding, involves creating two separate files: one linking an individual's identifying information to pseudonyms (substitute values) and another containing only the pseudonyms along with non-identifying research data. The former file is known as a key file or code list and must be kept separately from the latter file. See Figure 2 below for an example of pseudonymisation. Figure 2: Examples of pseudonymisation of names The key file should be securely stored and protected with measures such as password protection and encryption to prevent the re-identification of research participants. Encryption is the process of securing (personal) data by converting its storage location, such as a file, folder, USB device, hard disk drive, into a coded format, making it accessible only to those with the decryption key.
8 2.2.2 Anonymisation Anonymisation is another de-identification process which involves permanently deleting direct and/or indirect identifiers from the data, such that there is no way to link back to individuals (research participants) and the research information they have supplied. In this case, opposed to pseudonymisation, there is no key file that would permit re-identification and this process is irreversible as all identifiers are removed. An example of anonymising names can be found in Figure 3. Figure 4 shows the comparison between pseudonymisation and anonymisation. Figure 3: Examples of anonymisation of names Figure 4: Comparison of pseudonymisation versus anonymisation
15 Figure 7: Examples of suppression in textual data Figure 8: Examples of masking textual data
16 For more inspiration on how to anonymise textual data, please consult this illustrative example provided by the Consortium of European Social Science Data Archives (CESSDA), which demonstrates the process of anonymising an interview transcript. 3.2 NUMERICAL DATA Numerical data consist of numbers and is commonly organised in a table with rows and columns. Rows represent observations or attributes of each data subject, while columns display values for various variables or attributes of interest to researchers. Generalisation Generalisation techniques reduce the granularity of the attributes/records in the data. It can be performed by: (1) Categorising personal identifiers. For example, age is replaced by an age category or range. Geographic units are formed by combining all ZIP codes with the same four initial digits, which contain more than 20,000 people. (2) Top and bottom coding of the upper or lower ranges. Very high and low values are grouped into categories to minimise identifiability due to outliers. For example, ages above or below certain limits are classified as groups to prevent the identification of exceptionally older or younger individuals. A top code of“X or more”could be applied to avoid identifying older subjects. (3) Collapsing and/or combining variables. Merging data recorded in two or more variables into a single category. This is particularly useful if the initial data collection creates several categories with very few data subjects in each. (4) Perturbation can be used where small changes in value are acceptable. This technique involves rounding, adding noise, or replacing actual values with simulation values. It should not be used when data accuracy is critical. For examples on how to generalise (categorise and perturbate numerical data), see Figure 9 (next page). Rounding Rounding returns a number rounded to a base number to reduce precision. It can be used to deal with personal data, such as age, height and weight. Various tools can be used to apply the rounding technique, such as Excel, R, or Python. For instance, it is possible to use the MROUND function in Excel to round numerical identifiers. MROUND function will round the original value to a nearest value that is estimated by dividing the base number that is selected as parameter. For example, the function will turn the age of 44 to 40 when the base number is 20. Other practical examples of rounding are demonstrated in Figure 10 (next page).
17 Figure 9: Examples of generalising numerical data Figure 10: Examples of rounding numerical data
18 Noise addition Noise addition involves adding or subtracting the original values with a random number. This can be done manually or by means of tools such as Excel, R, or Python. The technique can be used for a few values in a dataset, such as outliers, or for the entire dataset. Large random values result in higher levels of noise. As a practical example, adding noise to the identifier “age” with R involves the following two steps: (1) generate a noise parameter using the sample(1:n) function to create a random number between 1 and n; (2) add or subtract this noise value from the original age. For instance, if the original age is 44 and the noise parameter is 5, the modified age becomes 49. See Figure 11 for additional examples of noise addition. Note that the results of noise addition are irreversible because each modification introduces randomness to the original values. Furthermore, when handling data with one or more continuous variables, the distribution of continuous variables should be preserved to maintain more information in the data. Figure 11: Examples of noise addition for numerical data
19 Suppression Suppression involves deleting identifying information from the data. Some attributes (variables) or records (observations) can be removed in some instances to make the identifier no longer unique. This is useful, for example, in the case of outliers that can easily allow the re-identification of an individual. Swapping The swapping technique is also known as shuffling or permutation to alter the association between the participants and their attributes. For instance, consider a dataset with 100 participants labelled as Participant 1 through Participant 100. Each participant has attributes such as income, education level, and job position. To reduce the risk of reidentification, attributes can be swapped between participants, for example exchanging the income of Participant 1 with that of Participant 50 and shuffle the job title for Participant 1 and Participant 87. By swapping the income and job title of participants, it reduces the likelihood of identifying an individual through both attributes. This process can be applied across the entire dataset, which effectively breaks the direct link between each participant and their original set of attributes. In this way, the distribution of the data would not be changed even though each participant’s attributes are altered. Note: This technique should not be used when the data has been collected to study relationships between attributes of the same research participants, such as the association between a job position and education level for each participant in the previous example. Masking numerical data Masking numerical data involves entirely or partially replacing digits with special symbols (such as '*' or 'x') in the same way words are masked in textual data as described in section 3.1. For instance, it is possible to use the REPLACE function in Excel to mask numerical identifiers such as phone number, income, email, IP address, etc. The REPLACE function requires three inputs: (1) the cell containing the original value, identified by its column and row (e.g., A2); (2) the starting and ending position of the digits/ characters within all content stored in the cell to be replaced; and (3) the symbol used to replace the original digit(s)/ character(s), such as ‘*’ or ‘x’. As an example, to mask the first three digits of an income value '58435' in cell A2 in Excel, the following formula can be used: Replace (A2, 1,3, “*”), which returns '***35 '. Additional examples of masking are shown in Figure 12 (next page).
20 Figure 12: Examples of masking numerical data
21 3.3 AUDIO-VISUAL DATA Audio-visual data include various forms of multimedia content, such as video and audio recordings, including voices and visual components. In research, these types of data often capture interviews, workshops, focus groups, places, or any other form of communication with research participants involving their voices and/or images. This section outlines some of the techniques used to anonymise audio-visual data. Voice transformation/distortion This technique involves disguising voices such as altering the pitch in a recording. Keep in mind that transforming or distorting the voice can alter acoustic characteristics, which decreases the usefulness of the data. Therefore, the decision to modify sounds should be made based on your research objectives. Image blurring, pixelation or obscuring Video images or pictures of places or individuals can be blurred or pixilated to reduce the level of detail in an entire image or part of an image or picture, which makes it harder to identify a person, place or object, etc. This technique involves reducing the level of detail in entire or part of images, making it harder to identify for example a person, location, licence plate, private spaces etc. For instance, you can pixelate the face of a person. See Figure 13 for an example. Figure 13: Example of a pixelated face. (Derived from pexels. CC0 by Pixabay) Simple distortion of the audio or imagery can only prevent identification based on the speaker’s voice or image but it cannot remove other identifying information that might exist in the recording, such as speaking patterns or the mentioned personal information.
22 Muting or bleeping out the identifying information In audio or video recordings, you can also mute, remove or bleep out identifying information to anonymise the audio-visual data. For example, you can mute, remove or bleep out identifiers such as names, places, or other information. Please see Figure 14 below, where the personal name has been muted in the audio recording. Figure 14: Demonstration of removing a part of an audio recording as illustrated by the arrow Software that can be used for audio-visual data There are various software and tools that can be used to anonymise audio visual data. The following are some examples: • For audio data: audio editing software such as Audacity can be used to transform or distort the voice or bleep out personal information. • For video data: video editing software such as Adobe Premiere can be used to blur or pixelate video recordings or transform, distort or bleep out information in the audio recording. • For images: image editor tools such as Adobe Photoshop and Paint can be used to blur, pixelate or obscure faces or personal information in pictures and images.
23 3.4 GEOSPATIAL DATA This section presents examples of the techniques commonly used for de-identifying geospatial data. Geospatial data refer to information linked to specific geographic locations. Geospatial data are often used in research for analysis, visualisation, and understanding of relationships and patterns within a geographic context. Geospatial data anonymisation involves techniques that reduce spatial or temporal resolution, i.e., the precision or scale at which data is captured (e.g., exact GPS coordinates vs. general area), and granularity, i.e., the level of detail in data attributes (e.g., full address vs. city name). Commonly used techniques to anonymise geospatial data include spatial aggregation, attribute anonymisation, and spatial displacement. Spatial aggregation The aggregation technique helps mask the exact locations while preserving overall spatial patterns and trends. It can take two forms: The first form is known as area aggregation, which involves reducing the level of detail by summarising spatial details (the location of a household or a hospital) into larger spatial units, such as zip codes, cities, census blocks or any other large administration or geographic units. Figure 15: Spatial aggregation of individual cases using census administrative units Figure from “Ensuring confidentiality of geocoded health data: Assessing geographic masking strategies for individual‐level data,” 2014, by P.A. Zandbergen, Advances in medicine, 567049. CC-BY 3.0 The second form is point aggregation, which is assigning multiple individual records to one geographic coordinate. Examples include a population dot map where one dot represents 100 persons (see Figure 15).
24 Attributes anonymisation The technique involves deleting or modifying any attribute or variable that could identify or link to individuals. Attributes such as the coordinates of households, exact locations of rare diseases, cadastral boundaries, ages, income, crime incidences, etc. can be deleted or modified similarly to numerical or textual data, as described in the previous sections. Spatial displacement/adjusting spatial coordinates Spatial displacement is a technique that involves shifting spatial data points (coordinates) to a different location within a specific range or distance by, for example, adding or subtracting a fixed or random number to both the x-axis and the y-axis. Spatial displacement preserves the spatial distribution and density of your data while making it difficult to pinpoint the exact original location of each point. It involves altering data accuracy to weaken links between the data and the individuals. This would mean that a map feature, for example, a point, is displaced to a new location (d-distance) away from its original location. Figure 16: The original coordinates of the University of Twente on Google Maps (52.24001648578193, 6.849739376582042) are displaced to new random coordinates (52.245130, 6.837690) at a given distance