After weeks of recruiting participants and conducting interviews, you have a folder full of audio recordings. Before you can start coding and identifying themes, those recordings need to become text that you can read, search and analyse. This step, transcription and data preparation, is often underestimated. It takes time, involves many small decisions, and has important implications for the quality, ethics and trustworthiness of your analysis.
This post walks through the whole process: choosing a transcription approach, deciding on the level of detail, using software and AI tools responsibly, anonymising transcripts, checking accuracy, and organising your data so that analysis runs smoothly.
Why transcription matters
Transcription is not a purely technical task. When you turn speech into text, you make choices about what to include: pauses, laughter, repetitions, false starts, dialect and non-verbal cues. These choices shape what you can later see in your data. Many qualitative researchers describe transcription as the first stage of analysis, because careful listening helps you become familiar with the data and notice early ideas.
Poor transcription can lead to misquoted participants, lost meaning and weaker findings. Good transcription gives you a reliable foundation for coding and interpretation.
Step 1: Decide on the level of detail
There are several styles of transcription. The right one depends on your research question and analytic method.
Verbatim (full) transcription
Everything that is said is written down exactly, including fillers (“um”, “you know”), repetitions, false starts and incomplete sentences. Pauses, laughter and interruptions may also be noted. This is common in many qualitative approaches and is essential when the way people speak is part of your analysis.
Intelligent (clean) verbatim
The content is kept word for word, but fillers, stutters and false starts are removed to improve readability. Meaning must not be changed. This style is often used in thematic analysis where the focus is on what people say rather than how they say it.
Detailed or specialised transcription
Approaches such as conversation analysis or discourse analysis use detailed notation systems to record timing, overlaps, intonation and emphasis. If you are using these methods, follow the conventions of your chosen approach, such as the Jefferson transcription system used in conversation analysis.
Summary or partial transcription
Only relevant sections are transcribed, or the content is summarised. This saves time but risks losing important data and is usually harder to defend in a PhD unless it clearly fits your method.
Whichever style you choose, decide in advance, write down your transcription conventions, and apply them consistently across all interviews.
Step 2: Create transcription conventions
A simple conventions document helps you and any transcribers stay consistent. It might include:
- How to label speakers, for example “I:” for interviewer and “P07:” for participant 7.
- How to mark pauses, such as “(pause)” or “(3 sec)”.
- How to show non-verbal sounds: “[laughs]”, “[sighs]”.
- How to indicate inaudible sections: “[inaudible 12:34]” with a timestamp.
- How to show uncertain words: “[unclear: shift manager?]”.
- How to mark interruptions or overlapping speech.
- How to record identifying information that will later be anonymised.
- Whether to include timestamps at regular intervals.
Step 3: Choose how to transcribe
Transcribing yourself
Transcribing your own interviews takes time, often several hours for each hour of audio, depending on detail, audio quality and typing speed. However, it brings real benefits: you become deeply familiar with your data, you hear tone and emphasis that may be lost later, and you can make notes of early analytic ideas.
Useful tools include foot pedals, playback speed controls and keyboard shortcuts in transcription software.
Using a professional transcription service
This saves time but costs money. If you use a service:
- Check that your ethics approval and consent forms allow a third party to access recordings.
- Use a reputable service with a confidentiality agreement.
- Check where data are stored and whether this meets your institution's data protection requirements.
- Provide your transcription conventions.
- Check the transcripts carefully against the audio.
Using automatic speech recognition and AI tools
Automatic transcription tools have improved considerably and can produce a first draft quickly. Many video conferencing platforms also offer automatic transcripts. However:
- Accuracy varies with audio quality, accents, technical vocabulary and overlapping speech.
- Errors can change meaning in subtle ways, such as “can” versus “can't”.
- Uploading recordings to cloud-based services may conflict with your ethics approval or data protection rules.
If you use these tools, check your institution's policy, use approved platforms where possible, and always correct the transcript manually by listening to the full recording. Treat the automatic output as a draft, not a final transcript. Mention the tool and your checking process in your methodology.
Step 4: Check accuracy
Whoever produces the transcript, accuracy checking is essential. A common approach is to listen to the whole recording while reading the transcript and correcting errors. Pay special attention to:
- Names, technical terms and numbers.
- Negations and qualifiers, such as “not”, “never”, “sometimes”.
- Sections marked as inaudible or unclear.
- Speaker labels, especially in focus groups.
Some researchers also offer participants the chance to review their transcript (sometimes called transcript review or member checking). This can improve accuracy and respect participants' ownership of their words, but it also has drawbacks: participants may want to delete or change statements, and it adds time. Decide whether it fits your methodology and ethics approval.
Step 5: Anonymise or pseudonymise
Most qualitative studies promise participants confidentiality. Transcripts must therefore be de-identified before analysis and certainly before any quotes are shared.
- Replace names with pseudonyms or participant codes.
- Remove or generalise identifying details, such as specific job titles, workplace names, towns, dates or unusual events. For example, “the head of cardiology at City Hospital” might become “a senior clinician at a large hospital”.
- Be careful with combinations of details that together could identify someone, especially in small communities or organisations.
- Keep a secure key linking codes to identities, stored separately from the transcripts, if your ethics approval permits this.
Record your anonymisation decisions in a log so that changes are consistent across transcripts.
Step 6: Organise and store your data
Good organisation saves time later.
- Consistent file names: for example “P07_Interview_Transcript_v2_anon.docx”.
- Version control: keep the raw transcript, the checked transcript and the anonymised transcript as separate, clearly labelled versions.
- Secure storage: use institution-approved storage, with encryption and password protection as required by your data management plan.
- Participant information sheet: create a summary table with participant codes and key characteristics (such as role, years of experience and interview length) for use in your methods chapter.
- Delete recordings according to your ethics approval, once transcripts are checked, if required.
Step 7: Prepare transcripts for analysis
Before coding, format transcripts to make analysis easier:
- Use a clear font, line numbers or paragraph numbers for referencing quotes.
- Add a header with participant code, date, interview length and interviewer.
- Leave a wide margin if you plan to code on paper.
- If using qualitative data analysis software such as NVivo, ATLAS.ti or MAXQDA, check the recommended file format and heading styles for importing.
- Attach relevant memos or field notes written after each interview.
Step 8: Begin familiarisation
Once transcripts are ready, read them in full, ideally while listening to the audio at least once. Write brief memos about your first impressions, interesting ideas and questions. This familiarisation step is the bridge between data preparation and formal coding.
Transcribing interviews conducted in another language
Many researchers conduct interviews in a language different from the language of their thesis. For example, you may interview participants in Tamil, Malay or Mandarin but write your thesis in English. This adds an extra layer of decisions.
- Transcribe first, then translate. A common approach is to transcribe in the original language and then translate the transcripts. This keeps an accurate record of what was actually said and allows you to return to the original wording during analysis.
- Analyse in the original language where possible. If you are fluent, coding in the original language can preserve nuances that are lost in translation. You can then translate only the quotes used in the thesis.
- Handle culturally specific terms carefully. Some words have no direct equivalent. Keep the original term in italics with a brief explanation, rather than forcing an imperfect translation.
- Check translations. Ask a second bilingual person to review a sample of translated passages, especially the quotes you plan to publish.
- Be transparent. Explain in your methodology who transcribed and translated the data, their language background, and how translation quality was checked.
If you use a professional translator, the same confidentiality and data protection rules apply as for transcription services. Make sure their involvement is covered by your ethics approval.
Time planning
Students often underestimate how long transcription and preparation take. Build realistic time into your research plan for:
- Transcribing or correcting automatic transcripts.
- Accuracy checks.
- Anonymisation.
- Formatting and importing into software.
Transcribing interviews soon after conducting them, rather than all at the end, spreads the workload and supports ongoing analysis, which also helps you judge saturation.
How to report transcription in your thesis
Include in your methodology chapter:
- Who transcribed the interviews and how.
- The transcription style and conventions used.
- Any software or automatic transcription tools used, and how output was checked.
- How accuracy was verified.
- Anonymisation procedures.
- Data storage and security measures.
Common mistakes to avoid
- Using automatic transcripts without careful checking.
- Uploading recordings to unapproved online services.
- Inconsistent conventions across transcripts.
- Incomplete anonymisation, especially of indirect identifiers.
- Waiting until all interviews are finished before transcribing any.
- “Tidying” participants' language so much that meaning changes.
Final thoughts
Transcription and data preparation are the foundation of qualitative analysis. Choose a transcription style that fits your method, write clear conventions, use tools responsibly, check accuracy carefully, anonymise thoroughly and organise your files well. The time you invest here will make coding smoother, protect your participants and strengthen the credibility of your findings.
Do you transcribe your interviews yourself or use a tool? Share your experience in the comments.