Anonymising interview transcripts
Anonymising looks like a chore at the end of a project and is in fact a methodological decision with consequences for the analysis. This guide shows what genuinely identifies people, what a properly anonymised passage looks like, and why expert interviews need their own approach.
Two terms that do not mean the same thing
Pseudonymisation replaces names with codes such as P3. Through a separately held key file the link remains restorable. Such data is still personal data and remains fully within the scope of the GDPR.
Anonymisation removes the link so far that attribution is no longer possible with additional knowledge and reasonable effort. Only then does the GDPR stop applying, and only then can a transcript be published without further conditions.
The workable route for most projects combines both: analyse pseudonymised, keep the key file separate and encrypted, and produce an anonymised version for the appendix.
What actually identifies people
The name is the most obvious feature and rarely the most dangerous. What identifies is the combination. Four details that look harmless alone often pin down exactly one person.
| Feature type | Example in the raw transcript | Anonymised version |
|---|---|---|
| Direct name | “Sarah brought that in back then” | “the then head of nursing brought that in” |
| Organisation | “here at St Mary's” | “at a large teaching hospital” |
| Place and region | “in the Peak District” | “in a rural district” |
| Rare role | “as the only stroke unit coordinator on site” | “in a coordinating role on site” |
| Datable event | “since the 2019 refurbishment” | “since a major refurbishment a few years ago” |
| Organisational figures | “there are only twelve of us” | “we are a small team” |
The last row is the one most often missed. Team sizes, bed numbers and turnover figures narrow the field of possible organisations drastically.
One passage, worked through
“So when I joined Hartley Engineering in Sheffield in 2019, I was the first woman in production management there at all. We had 40 people on the floor back then, and my boss, Mr Aldridge, watched that very closely.”
Six identifying features, highlighted.
“So when I joined [company] in Sheffield in 2019, I was the first woman in production management there at all. We had 40 people on the floor back then, and my boss, [name], watched that very closely.”
Four features remain. City, year, sector, headcount and the distinction of being the first woman in that role make the person identifiable in a city of that size.
“So when I joined a mid-sized manufacturing firm in the north of England a few years ago, I was the first woman in a production leadership role there. The department had a few dozen staff at the time, and the management of the day watched that very closely.”
The analytically relevant core survives intact: first woman in a leadership role, mid-sized manufacturing, attentive management.
The expert interview problem
Here a genuine conflict arises. The evidential weight of an expert interview rests on the interviewee having particular knowledge. Take the position away through anonymisation and the statement loses its footing. Leave it in and the person is often identifiable.
Three workable routes, depending on how sensitive the topic is:
- Role description instead of name. The function is described coarsely enough that several people fit it. Sound where the population is large enough.
- Attribution with express agreement. The person approves the quotation and consents to being named. Common and methodologically clean in interviews with industry or association representatives, as long as the consent covers it.
- Tiered versions. The thesis names roles, the submitted appendix is fully anonymised, and only the supervisor holds the mapping.
What does not work is deferring the decision. It belongs in the consent form, because the interviewee has to know how they will appear. A template is at Interview consent form.
How to do it in practice
- Fix the rules first. Which feature types get replaced how, written down in half a page. That note moves into your methods section.
- Use consistent placeholders, such as [organisation A], [city 1], [person 3]. Inconsistent replacements make cross-case comparison painful later.
- Anonymise in the first pass, alongside the correction listening. Two passes is one too many.
- Keep the key file separately, encrypted, not in the same folder as the transcripts.
- Have someone else read it. A person who does not know the case reads the anonymised version and says what stands out. That reliably finds what you read past.
- Delete the raw recordings once the purpose is met.
Note: this article is a practical overview, not legal advice.
Why a clean transcript shortens the job
Anonymising goes faster when it is clear who is speaking and when. With speaker changes marked, you can tell instantly whether a name came from the participant or from you. With timestamps set, you can re-listen at any unclear point instead of searching the recording.
Nodl delivers both from the uploaded audio: separated, colour-coded speaker passages and a timestamp per passage that jumps straight into the audio. Processing and storage take place in Germany, and content is not used to train models.
Common questions
In small language communities a marked accent can be identifying, especially combined with occupational detail. Content analysis usually standardises into written form anyway, which solves the problem in passing. Where speech form is the object of study, the trade-off belongs in the methods section.
Enough that the remaining description fits a sufficient number of people. A usable test: could somebody in that sector read this and say who it is? If the answer is maybe, you have not replaced enough.
Barely, if you replace features rather than strike them. St Mary's becomes a large teaching hospital, and for the analysis it was the type of institution that mattered, not the name. Value is only lost when whole passages disappear.
Once it is no longer needed, yes. While it exists the transcripts are pseudonymised, not anonymised. Deleting it is precisely the step that turns pseudonymised material into anonymous material.