How to transcribe an interview
Transcription is the part of qualitative research nobody budgets for and which eats the most time anyway. This guide covers which conventions you actually need, how much work is realistic, and where automation helps without damaging the quality of your analysis.
What a transcript has to deliver
A transcript is not an end in itself; it is the working basis for analysis. Three requirements follow. It has to be complete, so nothing relevant is lost. It has to follow traceable conventions, so quotations hold up. And it has to make findable where in the audio a statement sits, so contested passages can be checked.
Anything beyond that is effort without return. Noting every pause to the tenth of a second for a content analysis produces data you will never use.
The realistic time cost
5–10×
the audio length if you type it yourself
40–80 h
for eight 60-minute interviews
1 h
correction per hour of audio with automatic transcription
The rule of thumb for manual transcription is five to ten times the audio length, depending on recording quality, speech rate and typing speed. With detailed conventions the multiplier climbs considerably further.
| Interview length | manual | automatic plus correction |
|---|---|---|
| 30 minutes | 2.5 to 5 hours | about 30 minutes |
| 60 minutes | 5 to 10 hours | 30 to 60 minutes |
| 8 interviews of 60 minutes | 40 to 80 hours | 4 to 8 hours |
The third row is why the tooling question gets asked at all. Eight interviews is a common scale for a master's thesis, and 40 to 80 hours of typing is one to two full working weeks.
Which conventions you need
There is no universal standard. Several systems are established, at different levels of detail.
Intelligent verbatim
The default for content analysis. Speech is smoothed into standard written form, sentence structure is preserved, and fillers without meaning are dropped. Repetitions and stumbles are cleaned up unless they carry content.
Typical decisions:
- Speaker labels such as I for interviewer and P for participant
- Inaudible passages marked (inaudible), guesses placed in brackets
- Non-verbal events such as laughter noted only where they change the meaning
- Pauses from roughly three seconds marked as (pause)
- Timestamps at each speaker change or at fixed intervals
Detailed systems
Jefferson notation and comparable systems capture emphasis, overlaps, elongation and pause length precisely. They exist for conversation and discourse analysis, where the how of speaking is under study. For a content analysis they are oversized.
The practical test: if your research question would still be answerable had the person written the content down instead of speaking it, intelligent verbatim is enough.
Accents and dialect
With strongly accented recordings you face a decision: render into standard written form, or transcribe close to the sound. For content analysis, standardising is usual, with a note in the methods section. Automatic recognition degrades noticeably here, so allow more correction time than for standard speech.
More important than the choice itself is that it is made once and applied identically across every interview. Inconsistent transcripts are more vulnerable in a viva than a well-argued simple approach.
The workflow in six steps
- Fix your conventions before the first interview is transcribed. Half a page of your own decisions is enough and later moves into the methods section.
- Back up the recording. A copy in a second location before anything else happens. Lost interview recordings cannot be repeated.
- Produce the raw transcript, automatically or by hand.
- Correction pass with the audio. Listen and read through once, completely. This step is not optional, because automatic recognition reliably mangles technical terms, proper nouns and negations. A missed “not” inverts a finding.
- Anonymise. Replace names, places and organisations, including those of third parties.
- Format and export into whatever your analysis software expects.
Step four is the one most often cut short, and the one that determines the quality of the whole project.
Tools and what they are good for
| Approach | Strength | Limit |
|---|---|---|
| Manual with a playback aid | Free, full control, no data leaves your machine | The entire time cost stays with you |
| Local models on your own computer | No transfer to third parties, so no processing agreement | Setup effort, compute time, often no speaker separation |
| Online transcription services | Fast, speaker separation, timestamps | Data protection has to be checked, processing location varies |
| Analysis software with a transcription module | Transcript and coding in one tool | Licence cost, learning curve |
The choice depends less on features than on two questions: how many hours of audio do you have, and what does your institution require on data protection. For two short interviews a tool change is not worth it. From about five hours of material, automation always pays.
Data protection has its own guide, because in practice it decides whether a tool is permitted at all: Interview transcription and the GDPR.
From transcript to analysis
The transcript is an intermediate state, not a result. Content analysis continues with coding, assigning passages to categories. Whether categories come from theory or are developed from the material depends on the approach you chose.
That step benefits from what you built into the transcript. Clean speaker labelling lets you attribute statements without doubt. Timestamps make re-listening fast. Consistent conventions across interviews are what make cross-case comparison possible at all.
How Nodl produces the raw text
Nodl takes the recording, produces the transcript and, if you want, a structured document in a format you define. Three properties matter for interviews:
- Existing recordings can be uploaded, not only spoken in live
- With several voices the passages are separated and colour-coded by speaker
- Each passage carries a timestamp that jumps straight to that point in the audio
Recordings are stored encrypted on servers in Germany, the language models run inside the EU, and content is not used for training. A single recording may be up to one hour long; longer interviews have to be split first.
None of this replaces the correction pass, and no automatic transcription does. It shifts the work from ten hours of typing to one hour of checking.
Common questions
Your examination board or supervisor decides, and anonymised transcripts in a digital appendix are commonly required. Clarify early, because it determines how thoroughly you must anonymise. An appendix that gets published needs genuine anonymisation, not just coded names.
Usually by speaker code plus a locator, either a line number or a timestamp. That is why exporting the transcript with continuous line numbering is worth the trouble. The required form is in your department's guidance.
Prevention is the only good answer: record close to the microphone, avoid rooms with hard surfaces, and test for thirty seconds before the interview starts. For recordings you already have, only more correction time helps, and unclear passages get marked (inaudible) rather than guessed.
Yes, it is common and methodologically unproblematic as long as you do the correction pass yourself. In data protection terms a transcription agency is a processor, so you need an agreement under Article 28 GDPR and a mention in the consent form.