How to handle personally identifiable information in research recordings and transcripts while maintaining analysis quality and GDPR compliance
User research generates some of the richest qualitative data a product team can work with. It also generates some of the most sensitive. A single interview recording can contain a participant's full name, workplace, health details, financial situation, and opinions about their employer—all in the space of a few minutes.
If your organization operates under GDPR or collects data from EU residents, how you store, process, and share those recordings and transcripts is not just a best practice question. It is a legal one.
This guide walks through a practical approach to handling personally identifiable information (PII) in research data. The aim is to help you protect participants and stay compliant without stripping your data of the context that makes it useful for analysis.
Why PII in research recordings needs special attention
Quantitative data is relatively straightforward to anonymize. You can strip identifiers from a spreadsheet and still analyze the numbers. Qualitative research data is different. Participants tell stories, describe personal experiences, and mention specific people, companies, and places. PII is woven into the substance of what they are saying.
This creates a tension. If you aggressively redact everything, you risk removing the context that makes qualitative data valuable. If you leave everything intact, you risk exposing participants to harm and your organization to regulatory penalties.
GDPR does not prohibit research. It requires that personal data be processed lawfully, transparently, and with purpose limitation. Under Article 5, data must be adequate, relevant, and limited to what is necessary. Under Article 89, research is given certain exemptions—but those exemptions still require appropriate safeguards, including pseudonymization where feasible.
The practical challenge is implementing those safeguards across recordings, transcripts, highlight reels, and shared analysis—without creating so much friction that researchers stop doing their jobs well.
Understanding what qualifies as PII in research contexts
Before you can protect PII, you need to know what you are looking for. In a research recording, PII falls into several categories.
Direct identifiers
These are data points that identify a person on their own:
- Full name
- Email address or phone number
- Home or work address
- Government-issued ID numbers
- Photos or video of a participant's face
Indirect identifiers
These are data points that identify a person when combined with other information:
- Job title and company name (especially in small organizations)
- Age, gender, and geographic location mentioned together
- Descriptions of unique personal situations ("I'm the only accessibility specialist at my company")
- Names of family members, colleagues, or managers
- Details visible in a video background—mail, badges, calendar screens
Special category data
GDPR Article 9 defines special categories of personal data that require additional protections. These include health information, racial or ethnic origin, political opinions, religious beliefs, and trade union membership. Research participants frequently volunteer this kind of information unprompted, particularly in interviews about healthcare products, financial services, or workplace experiences.
The point is that PII in qualitative research is rarely limited to a name and email address. Researchers need to develop a habit of recognizing PII in all its forms—spoken, visual, and contextual.
Building a PII handling workflow for research data
Treating PII protection as an afterthought—something you deal with when someone raises a concern—creates risk. A better approach is to build PII handling into your research workflow from the start.
Before the session: set the ground rules
Define your lawful basis for processing. Under GDPR, consent is one lawful basis, but it is not the only one. Legitimate interest (Article 6(1)(f)) can also apply, particularly for B2B research. Work with your legal or data protection team to determine which basis applies and document it.
Write a clear participant information sheet. This should explain what data you will collect, why, how it will be stored, who will have access, how long it will be retained, and how participants can exercise their rights (access, rectification, erasure, withdrawal of consent). Avoid legalese. Participants should genuinely understand what they are agreeing to.
Design your consent process. If consent is your lawful basis, it must be freely given, specific, informed, and unambiguous. Use separate consent checkboxes for recording audio, recording video, and sharing clips with internal teams. Do not bundle these together.
Set up storage with access controls. Decide where recordings and transcripts will be stored before you create them. Ensure the storage location has role-based access controls, encryption at rest and in transit, and audit logging. Cloud-based research platforms like Dovetail offer these capabilities natively, which reduces the operational burden compared to managing files on shared drives.
During the session: minimize unnecessary PII collection
You cannot always prevent participants from sharing PII during a session—nor should you interrupt their flow to stop them. But you can take steps to reduce unnecessary collection:
- Ask participants to avoid sharing their full name on camera if the session is being recorded. Use first names or participant codes.
- If you are conducting a remote session, ask participants to use a virtual background or check that their physical background does not reveal identifying information.
- If the research does not require video, consider recording audio only. This eliminates an entire category of visual PII.
- If screen sharing is involved, ask participants to close email, messaging apps, and browser tabs that might display personal information.
These steps are not foolproof, but they reduce the volume of PII you need to handle downstream.
After the session: redact, pseudonymize, and organize
This is where the most significant work happens. Once a recording exists, you need a consistent process for handling the PII it contains.
Transcription. If you are using automated transcription, review the transcript for PII before sharing it. Automated tools will faithfully transcribe names, email addresses, and other identifiers spoken aloud. Replace these with pseudonyms or placeholder labels ("[Participant 7]", "[Company A]").
Audio and video redaction. For recordings you plan to retain or share, redact PII from the media itself. This means bleeping or silencing names and identifiers in audio, and blurring faces, screens, or background details in video. Several tools exist for this, and some research platforms are beginning to offer automated PII detection and redaction.
Create a redaction key. Maintain a separate, access-controlled document that maps pseudonyms back to real identities. This allows you to re-identify participants if needed (for example, to honor an erasure request) without exposing their identities in day-to-day analysis files.
Tag and classify sensitivity. Not all sessions contain the same level of PII. A usability test where a participant navigates a prototype may contain minimal PII. A deep interview about financial hardship contains a great deal. Classify your recordings by sensitivity level and apply access controls accordingly.
Maintaining analysis quality after redaction
A common concern among researchers is that redaction will undermine the quality of their analysis. If you strip too much context, the data loses its richness. Quotes become vague. Themes become harder to trace back to specific participant experiences.
This concern is valid, but it is manageable with a few practices.
Use meaningful pseudonyms
Rather than replacing all identifiers with generic labels, create pseudonyms that preserve relevant context. Instead of "[Company A]," you might use "[mid-size fintech company]" if the company's industry and size are analytically relevant but its specific identity is not. Instead of "[Participant's manager]," you might use "[direct supervisor, engineering]."
The goal is to preserve the attributes that matter for analysis while removing the attributes that enable identification.
Preserve emotional and behavioral context
Redaction should focus on identifiers, not on the content of what participants said. A participant's description of their frustration with a product, the workarounds they have developed, or the language they use to describe a problem—these are the analytically valuable elements, and they rarely constitute PII on their own.
When you redact, keep the surrounding context intact. If a participant says, "My manager, Sarah, never checks the dashboard," the redacted version should read, "My manager, [name removed], never checks the dashboard"—not just "[redacted]."
Work from transcripts for day-to-day analysis
For most analytical tasks—coding, theming, pattern identification—transcripts are more practical than raw recordings. A well-prepared, pseudonymized transcript gives researchers everything they need to do rigorous qualitative analysis without repeated exposure to raw PII.
Reserve access to original recordings for situations where tone of voice, facial expression, or screen interaction genuinely adds analytical value. This reduces PII exposure across the team while keeping the full-fidelity data available when it matters.
Dovetail's approach to research data management supports this workflow. Transcripts, tags, and highlights can be organized and analyzed within the platform while access to underlying recordings is governed by permissions—making it possible to share insights broadly without sharing raw PII broadly.
GDPR compliance: key obligations for research teams
Beyond the practical workflow, there are specific GDPR requirements that research teams need to address.
Data minimization
Collect only the personal data you need for your stated research purpose. If you do not need video, do not record video. If you do not need to know a participant's employer, do not ask. Review your screening questionnaires and discussion guides for unnecessary PII collection.
Purpose limitation
Personal data collected for research should be used for research. If you want to use a participant's quote in marketing materials or share their recording with a client, that is a different purpose and requires separate consent.
Storage limitation
Define how long you will retain recordings and transcripts, and delete them when that period expires. "We might need it someday" is not a compliant retention policy. Common retention periods for research data range from 6 to 24 months, depending on the nature of the research and organizational policy.
Right to erasure
Participants have the right to request deletion of their personal data. Your workflow needs to account for this. Can you locate all copies of a specific participant's data—recordings, transcripts, highlight clips, notes, coded excerpts—and delete them? If your data is scattered across shared drives, Slack messages, and presentation decks, this becomes extremely difficult.
Centralizing research data in a single platform significantly simplifies erasure requests. When all of a participant's data lives in one place, finding and deleting it is a straightforward operation rather than a scavenger hunt.
Data protection impact assessment
If your research involves large-scale processing of special category data or systematic monitoring of individuals, GDPR Article 35 may require you to conduct a Data Protection Impact Assessment (DPIA). Consult your Data Protection Officer or legal team to determine whether this applies to your research program.
Common mistakes to avoid
Storing recordings on personal devices. Researchers sometimes download recordings to their laptops for convenience. This creates uncontrolled copies of PII outside your security perimeter. Use cloud storage with proper access controls instead.
Sharing unredacted clips in Slack or email. A compelling participant quote or video moment is easy to share casually. But every share creates a new copy of PII in a location you may not control. Share clips through your research platform, where access is governed and auditable.
Treating consent as a one-time event. Consent is ongoing. If you change how you use research data—for example, by sharing it with a new team or using it for a new purpose—you may need to re-consent participants.
Assuming anonymization is permanent. True anonymization (where re-identification is impossible) is very difficult to achieve with qualitative data. A participant's speech patterns, described experiences, and contextual details can enable re-identification even when direct identifiers are removed. Treat your data as pseudonymized rather than anonymized, and apply protections accordingly.
Neglecting deletion schedules. It is easy to accumulate years of research recordings with no plan for deletion. Set calendar reminders or use platform features that flag data approaching the end of its retention period.
Building a PII-conscious research culture
Compliance is not just about tools and processes. It is about building habits across your research team—and across the product managers, designers, and stakeholders who consume research outputs.
Train everyone who handles research data on what PII looks like in qualitative contexts. Make redaction part of your standard operating procedure, not an optional extra. Include PII handling in your research onboarding for new team members.
Document your policies and make them accessible. A short, clear internal guide on "how we handle participant data" is more useful than a 40-page policy document that no one reads.
Finally, choose tools that make compliance easier rather than harder. Platforms designed for research data management—like Dovetail—can automate parts of the PII handling workflow, enforce access controls, support retention policies, and centralize data so that participant rights requests are manageable. The right infrastructure turns compliance from a burden into a default.
Conclusion
Handling PII in research recordings and transcripts is not a problem you can solve once and forget. It is an ongoing practice that requires clear policies, consistent workflows, and a team culture that takes participant privacy seriously.
The good news is that protecting PII and maintaining analysis quality are not opposing goals. With thoughtful pseudonymization, structured redaction, and centralized data management, you can preserve the richness of qualitative research while meeting your obligations under GDPR.
Start with your next study. Review your consent forms, set up your storage, plan your redaction workflow, and define your retention period. Each of these steps is straightforward on its own. Together, they form a system that protects participants, reduces organizational risk, and keeps your research insights intact.
FAQs
What counts as personally identifiable information in a research recording?
PII in research recordings includes any data that can directly or indirectly identify a participant. Obvious examples are names, email addresses, phone numbers, and physical addresses spoken aloud or visible on screen during a session. Less obvious examples include employer names, job titles in small organizations, descriptions of unique personal circumstances, or background details visible in a video feed—like mail on a desk or a name badge. Under GDPR, even pseudonymous data can qualify as personal data if it can be linked back to an individual using additional information you hold.
Can I keep full unredacted recordings if participants have given consent?
Consent alone does not eliminate your obligations under GDPR. Even with explicit consent, you must still adhere to data minimization principles—meaning you should only retain personal data that is necessary for the stated research purpose. You also need to define and enforce a retention period, ensure the data is stored securely with appropriate access controls, and honor participants' right to withdraw consent at any time. In practice, this means you should still plan for redaction or pseudonymization of data you do not need, even when you have consent, to reduce risk and demonstrate compliance.
How do I redact PII from recordings without losing research context?
The goal of redaction is to remove identifiers while preserving the meaning and emotional texture of what participants said. For audio, you can replace PII with a brief tone or silence and add a transcript annotation like '[name removed]' so analysts understand the gap. For video, blurring faces or on-screen details achieves a similar effect. To preserve context, create a redaction log that maps anonymized labels (e.g., 'Participant 7' or 'Company A') to their original values, and store that log separately under strict access controls. This way, analysts can follow the narrative without being exposed to raw PII, and you can re-identify if legally required.