An interview transcript records who said what, when, and which hesitations you keep. That last part decides everything else. Here is one sentence from a research interview, written down twice. In full verbatim: P03: I, um, I didn't (1.5) I didn't want to go back. In clean verbatim: P03: I didn't want to go back. Both are correct. They serve different readers, and the choice between them shapes the header, the tags, the hours you budget and what a co-author can check later. Below is that passage in both styles, three more transcripts for hiring, journalism and customer research, and a template you can copy. Afterwards you can pick a style and defend it. You can budget the hours with numbers you can cite, and hand a transcript to an ethics board without exposing a participant.
Key takeaways
- At the 164 to 196 words per minute measured across 2,438 recorded phone conversations (Yuan, Liberman and Cieri, 2006), a one-hour interview holds roughly 9,800 to 11,800 spoken words.
- Plan 5 to 10 hours of work per hour of interview audio typed by hand, proofreading included. Dresing and Pehl’s transcription handbook (2024) gives that range.
- Full verbatim keeps every um, false start and timed pause, written as (1.5) for 1.5 seconds. Clean verbatim removes them without changing a word. Pick one before you start.
- A research transcript has 4 parts: a header, speaker labels, timestamps and a tag legend. Timestamps take you back to the audio, while quotations cite paragraph numbers instead.
- One speech model scored 2.7 percent word error on clean read speech and 36.4 percent on a distant meeting microphone (Radford et al., 2022). Check every AI draft against the audio.
- A transcript that says P07 and keeps the name key elsewhere is pseudonymised, not anonymous, under Art. 4(5) GDPR. Store that key apart from the transcripts, always.
- Key takeaways
- Interview transcription example: one passage, two styles
- What goes into an interview transcript
- Interview transcript examples for three situations
- Verbatim, clean or edited
- Copy this interview transcript template
- How to transcribe an interview in seven steps
- How long an interview transcript takes
- Consent, anonymisation and GDPR
- Multilingual interviews and dissertations
- Interview transcripts with alugha
- Frequently asked questions about interview transcription examples
- Getting started
Interview transcription example: one passage, two styles
Start with the transcript itself. The passage below comes from a qualitative interview about returning to work after a long illness: three minutes, two speakers, the interviewer as INT and the participant as P03. Every participant, city and employer in the examples on this page is invented; the conventions are not.
Interview transcription example in full verbatim
Full verbatim keeps everything the recorder caught: fillers, repeated words, a word broken off halfway, every pause in seconds, and the interviewer’s small sounds of agreement. The header comes first, because a transcript without one can only be checked by the person who typed it.
STUDY RTW-2026, return to work after long illness
PARTICIPANT P03
INTERVIEWER INT
DATE 2026-03-14, video call, audio recorded
FILE RTW-2026_P03_2026-03-14.m4a
STYLE full verbatim
TIMESTAMPS at every speaker turn
CONSENT form C-P03, signed 2026-03-10
TAGS (seconds) pause, / broken word, [overlap], [laughs],
[inaudible hh:mm:ss], [CITY] [EMPLOYER] replacements
[00:03:12] INT: And the first day back. What do you remember?
[00:03:20] P03: I, um, I didn't (1.5) I didn't want to go back.
The badge, the, the badge didn't work. (1.0) I stood at
the gate at [EMPLOYER] and it just, uh, it just didn't
work. [laughs]
[00:03:34] INT: Mm-hm.
[00:03:35] P03: And I thought, okay, that's, that's a sign. I was
wo/ I was working from home for eleven months before
that, in [CITY], and, uh, [overlap] the whole team had
[00:03:47] INT: [overlap] Eleven months.
[00:03:48] P03: Yeah. Eleven. (2.0) The whole team had changed.
Nobody at the gate knew me. My manager was new, the,
uh, the desk was gone. Somebody had put a plant on it.
[00:04:02] INT: How did that feel?
[00:04:04] P03: Honestly? Like being a visitor. Like, you know, you
get the sticker and you, um, you wait. And I
[inaudible 00:04:12] for about an hour, I think.
[00:04:16] INT: What helped, in the end?
[00:04:18] P03: The phased thing. The (1.0) the four hours a day for
the first month. Without that I, I would have quit.Read it the way a co-author would. The (1.5) before the second “I didn’t” is a hesitation you can measure, and the repeat after it belongs to the answer. The “wo/” at 00:03:35 is a word broken off and started again. At 00:03:47 the interviewer talks over the participant, so both turns carry an [overlap] tag. The inaudible stretch keeps its time, and anyone can find the gap in the recording. None of this is decoration. If your question is how people talk about coming back, these marks are the evidence.
The same passage in clean verbatim
Clean verbatim runs a style pass over the same words. The fillers go, the stutters go, and so do the false starts and the interviewer’s “mm-hm”. Nothing is rephrased. The header changes in two lines: the style, and the tags that no longer occur.
STUDY RTW-2026, return to work after long illness
PARTICIPANT P03
INTERVIEWER INT
DATE 2026-03-14, video call, audio recorded
FILE RTW-2026_P03_2026-03-14.m4a
STYLE clean verbatim
TIMESTAMPS at every speaker turn
CONSENT form C-P03, signed 2026-03-10
TAGS [laughs], [inaudible hh:mm:ss], [CITY] [EMPLOYER]
[00:03:12] INT: And the first day back. What do you remember?
[00:03:20] P03: I didn't want to go back. The badge didn't work. I
stood at the gate at [EMPLOYER] and it just didn't
work. [laughs] And I thought, okay, that's a sign. I was
working from home for eleven months before that, in
[CITY], and the whole team had changed.
[00:03:47] INT: Eleven months.
[00:03:48] P03: Yeah. Eleven. Nobody at the gate knew me. My manager
was new, the desk was gone. Somebody had put a plant
on it.
[00:04:02] INT: How did that feel?
[00:04:04] P03: Honestly? Like being a visitor. You get the sticker
and you wait. And I [inaudible 00:04:12] for about an
hour, I think.
[00:04:16] INT: What helped, in the end?
[00:04:18] P03: The phased thing. The four hours a day for the first
month. Without that I would have quit.Count the words and the difference shows. P03 speaks 134 words in the full version and 106 in the clean one, about a fifth fewer. What the clean version loses is the (1.5) before “I didn’t want to go back”. A reader who wants to know what happened on the first day loses nothing. A reader who studies how hard that sentence was to say loses the one mark that showed it. That is why the style goes into the header, and why it stays the same in every file of the study.
What changed, and what did not
- Removed: “um” and “uh”, repeated words, the broken “wo/”, the timed pauses, the overlap marks and the interviewer’s backchannel. Two P03 turns merged once the “mm-hm” between them went.
- Kept: every word of substance in the speaker’s order, the laugh that tells you the badge story is told with irony, the inaudible marker with its time, and the replacements for the city and the employer.
- Never touched: the meaning, and the speaker’s own grammar.
A transcript is not the recording in text form. It is a set of decisions about what to keep.
The linguist Mary Bucholtz made the point in 2000: every transcript embeds interpretive and representational decisions. So name yours in the header, where the next reader will find it.
What goes into an interview transcript
Start with the parts. McLellan, MacQueen and Neidig wrote in 2003 that there is “no universal transcription format” that suits every qualitative approach, setting or framework. Every good transcript still has the same four parts.
The header block
The header answers the questions a second reader asks before the first line. It holds the study ID, participant code, ISO date, interviewer, mode, the recording’s file name, the style, the timestamp rule, a consent reference and the tags in use. Indeed’s example for hiring managers gets the header right, listing names, date, place and everyone present. Research keeps those fields and swaps the names for codes.
Two rules keep the header safe. It never holds the participant’s name, only the code, because the header travels with every copy of the file. And the consent line points to the signed form instead of repeating what the form says. The file name follows the same pattern of study, code and date, so it says nothing about the person either.
Speaker labels
Pick your labels once and keep them. Research uses INT for the interviewer and P01, P02 for participants. HR records often use initials, as Indeed’s example does. Full names belong in a published interview, with consent, or in an HR file with restricted access. If the file goes into analysis software, format matters: NVivo expects the speaker name followed by a colon and a space, and MAXQDA recognises several timestamp formats, including [hh:mm:ss.xx] and plain hh:mm:ss.
With two interviewers, number them: INT1 and INT2. In a group interview, give every participant a code and ask each person to say it once at the start, so the voices can be matched later. If you cannot tell who is speaking, write that down instead of guessing, for example P? for a participant you cannot identify. A wrong label does more harm than an honest question mark, because the analysis builds on it.
Timestamps
Set a timestamp at every speaker turn, or at a fixed interval such as every 60 seconds. It takes you back to the audio when a line reads oddly. Dresing and Pehl’s handbook (2024) is explicit that timestamps are not for citation. Quotations in a thesis cite paragraph numbers.
Write timestamps as hh:mm:ss in square brackets, counted from the start of the recording, not from the clock on the wall. If one interview is split across two files, restart at 00:00:00 in each and name the file in the header. Otherwise the times point nowhere.
A timestamp is how you find the audio again. It is not how you cite it.
Subtitle files work differently, because WebVTT and SRT carry a start and an end time for every cue by design. The video transcript examples show what those files look like.
Tags for what the words miss
Words are not the whole recording. A laugh, a gap, two people talking at once: each needs a tag, and each tag needs one meaning across every file of the study. List every tag you use in the header, so a co-author decodes the transcript the way you wrote it.
Symbols travel badly between traditions. In conversation analysis, (.) is a micropause under two tenths of a second. In the simple transcription rules of Dresing and Pehl’s German handbook, (.) is a pause of about one second, and // marks overlap. Neither system is wrong. A co-author who knows only one of them will still misread the other, unless the legend sits in the header. Describe, do not interpret: [laughs] is a tag, while [laughs nervously] is already a reading.

In short: a header, fixed speaker labels, timestamps and a tag legend. The exact format matters less than using the same one in every file of the study.
Interview transcript examples for three situations
Research is one use. Hiring, journalism and customer research each bend the rules, and two questions decide how: who reads the transcript, and whether how something was said counts, or only what.

Job interview transcript example
Indeed publishes a job interview example for hiring managers, and it earns its place. It has a proper header, initials as speaker labels and bracketed notes for what the words miss. It has no interval timestamps, though, and its budget of seven to ten hours for beginners comes without a source. The version below keeps Indeed’s shape and adds a timestamp at every turn and a consent line. A hiring record is applicant data, and the data protection rules below apply to it.
ROLE Senior accountant, finance team
CANDIDATE J.M.
PANEL R.K. (head of finance), S.T. (HR)
DATE 2026-05-06, 10:00, on site, room 2.14
FILE FIN-SA-2026_JM_2026-05-06.m4a
STYLE clean verbatim
TIMESTAMPS at every speaker turn
CONSENT Candidate agreed to recording and transcription for
the hiring file; deletion after the process.
[00:12:04] R.K.: Walk us through the month-end close you ran last.
[00:12:09] J.M.: We closed in six working days. I took over the
accruals and moved the reconciliations into one
shared checklist, so nobody waited for an email.
[00:12:24] R.K.: What went wrong the first time?
[00:12:27] J.M.: The intercompany balances. Two subsidiaries booked
the same invoice in different months. We found it on
day five, which was too late.
[00:12:41] S.T.: And what did you change?
[00:12:43] J.M.: A cut-off date both teams sign. It sounds small.
It saved us two days the next quarter.
[00:12:52] R.K.: How do you handle a disagreement with an auditor?
[00:12:56] J.M.: I ask which standard they are reading from, and
then I show my working. [laughs] Usually one of us has
the wrong version of the file.
[00:13:10] S.T.: Thank you. [phone rings] Sorry, one moment.Two details matter for hiring. Ask every candidate the same core questions, so the transcripts can be laid side by side and compared line by line. And let the consent line name the purpose and the end date. The GDPR asks for exactly that: under Art. 5(1)(e), personal data are kept no longer than the purpose needs, and a hiring file has a natural end. Keep tags such as [phone rings], too. An interruption can explain a short answer.
Expert interview edited for publication
Journalism edits more and alters less. The question is shortened, detours are cut, and the answer is trimmed to what the reader needs. Quoted words stay exactly as spoken. AP style does not alter quotations, not even to correct minor grammatical errors. The Oral History Association’s principles, adopted in 2018, add a step from oral history: whenever possible, the speaker gets the chance to review and approve the interview before it is used. Mark every cut with [...] and every word of yours with square brackets.
PIECE Company blog, "Running a plant through a heatwave"
SPEAKER A. Brandt, plant manager (named with consent)
INTERVIEWER Editorial team
RECORDED 2026-07-22, phone, 38 minutes
STYLE edited for publication, quotations unchanged
REVIEW Transcript approved by the speaker on 2026-07-29
[00:04:10] Q: What happens first when the temperature climbs?
[00:04:15] Brandt: We move the shifts. The early shift starts at
five, not six, so the heaviest work is done before
noon. [...] The machines cope better than people do.
[00:09:32] Q: Did production fall?
[00:09:35] Brandt: For two days, yes. After that [the plant] ran at
the same output as in May, because we had stopped
fighting the heat and started planning around it.
[00:21:48] Q: What would you tell another plant manager?
[00:21:51] Brandt: Ask your night shift. They knew in June what we
only understood in August.Read the marks. The […] at 00:04:15 is a cut. “[the plant]” stands in for the speaker’s “it”, which the reader could not resolve once the question before it was shortened. The timestamps stay in your working copy and come out of the published text. Keep them anyway: when the speaker asks what exactly they said, you find the passage in seconds. The review line in the header is your record that the speaker has seen the edit.
Customer interview with analysis codes
UX and market research code while they read. Braun and Clarke (2006) place transcription in the first phase of thematic analysis, familiarising yourself with the data, which includes noting initial ideas; systematic coding follows in phase two. So it pays to note first ideas and candidate codes beside the text on the first read. The example below writes them as // code: lines. They survive any editor and paste cleanly into a spreadsheet.
STUDY INV-2026, invoicing app, churn interviews
PARTICIPANT C07, finance lead at a 40-person agency
DATE 2026-04-02, video call
STYLE clean verbatim, codes added on first read
[00:06:40] INT: How do you send invoices today?
[00:06:44] C07: From the app, but I export every one to a PDF first
and check the VAT line by hand.
// code: workaround, trust in calculation
[00:06:58] INT: Why by hand?
[00:07:01] C07: Because in January it rounded wrong on two invoices
and a client noticed before we did.
// code: critical incident, reputational risk
[00:07:15] INT: What would make you stop checking?
[00:07:18] C07: Seeing the calculation. Not a promise, the actual
lines. And the renewal is hard to justify while I
still do this. [laughs]
// code: transparency request, price hesitationKeep code names short and reuse them. “Trust in calculation” in interview C07 has to read the same in C12, or the count at the end is wrong. First-read codes are candidates. In phase two they are merged, renamed or dropped, and a codebook fixes what each one means. The transcript text does not change when a code does, which is one more reason to keep codes on their own lines.
In short: hiring keeps a fair record, journalism keeps the quote exact, and research keeps its codes beside the words. The header tells the reader which of the three they are holding.
Verbatim, clean or edited
Three styles cover almost every interview. Rev, a large transcription service, gives the usual industry definitions. Verbatim means every word spoken is written down word for word. Clean verbatim is lightly edited for readability: stutters, fillers such as “um” and “like”, false starts and unintentional repetition come out, with “never any paraphrasing”. An edited transcript goes further and cuts whole passages for a reader outside the study. The names drift, too. Some vendors sell clean verbatim as “intelligent” or “smart” verbatim, and some edit a little more under those names. Pick one term, define it and use it in every header.

| Element | Full verbatim | Clean verbatim | Edited | |—|—|—|—| | Filler words (um, uh, like) | Keep | Remove | Remove | | Stutters and repeats | Keep | Remove | Remove | | False starts | Keep | Remove | Remove | | Interviewer backchannel (mm-hm) | Keep | Remove | Remove | | Laughter, sighs | Keep | Keep if it changes meaning | Keep if it changes meaning | | Timed pauses | Keep, in seconds | Remove | Remove | | The speaker’s grammar | Keep | Keep | Keep inside quotes | | Off-topic passages | Keep | Keep | Cut, marked […] |
When full verbatim is worth the hours
Full verbatim is slow. For some work it is also the only honest choice, because the way something was said is the data. Conversation analysis studies turn-taking, hesitation and overlap, and Gail Jefferson’s notation exists for exactly that. A left square bracket marks where an overlap begins. A period in round brackets marks a micropause under two tenths of a second, and a number in round brackets a timed silence. Hand transcription is the standard there. Legal records and oral history, where the exact words go on the record, sit close by.
It is not all or nothing, though. A study can transcribe every interview in clean verbatim and return to full verbatim, or to Jefferson notation, only for the passages it analyses closely. Note the switch in the header, with the time range. Then nobody reads the detailed passage as the standard of the whole study.
When clean verbatim is enough
Most thematic research does not need every um. Halcomb and Davidson argued in 2006 that verbatim transcription is not always necessary, especially in mixed-methods studies. Decide per study, not per interview. Mixed styles make transcripts look different for reasons that have nothing to do with the participants. And clean does not mean tidy: the speaker’s grammar, dialect words and unfinished thoughts that carry meaning stay as spoken. Whatever style you choose, check it against the audio. Blake Poland observed in 1995 that transcript accuracy is more often assumed than demonstrated. Hence the second listen in the workflow below.
In short: full verbatim when how something was said is the data, clean verbatim when what was said is the data, and edited only for a reader outside the study.
Copy this interview transcript template
Every example above follows one template. Copy it into any plain-text editor and delete the lines you do not need. Nothing in it depends on a particular tool.
STUDY / PROJECT <study ID and short title>
PARTICIPANT P01 (code only, never a name)
INTERVIEWER INT
DATE YYYY-MM-DD
MODE in person / phone / video call
FILE <study>_<code>_<date>.m4a
STYLE full verbatim / clean verbatim / edited
TIMESTAMPS every turn / every 60 seconds
CONSENT <form ID>, signed YYYY-MM-DD
TRANSCRIBED BY <initials>, YYYY-MM-DD
CHECKED BY <initials>, YYYY-MM-DD, against the audio
SPEAKERS INT = interviewer, P01 = participant
TAGS [inaudible hh:mm:ss] [word?] [overlap] (2.0)
[laughs] wo/ [CITY] [clarification]
[00:00:00] INT:
[00:00:00] P01:
[00:00:00] INT:
REPLACEMENTS logged in <file name>, stored separatelyThree rules for filling it in
Write every date as YYYY-MM-DD, so files sort by date in any folder and nobody confuses the 3rd of April with the 4th of March. Give each person one code across every file of the study, interviewers included if there are several. A participant interviewed twice keeps P03 both times, and the date tells the sessions apart. And if you will quote from the transcript, switch on paragraph numbering in your editor or analysis tool before you start. Those numbers are what your citations will point to, so they must not shift after the first quote. Once a transcript is checked, freeze it. Later corrections go into a new version with its own date in the file name.
In short: one header per file, one code per person, ISO dates, and paragraph numbers that never move.
How to transcribe an interview in seven steps
Seven steps, and the order matters. Anonymisation comes after the style pass, because replacing names in a text you are about to reformat means doing the job twice.

Record for the transcript
The transcript starts before the first question. Get consent on the recording itself, in the participant’s own words. Give each speaker a microphone if you can, and pick a quiet room over a convenient one. Switch off fans and notifications. On a video call, record each voice on its own track where the software allows it. Record ten seconds, play them back, and only then begin. A second device running as a backup saves the interview when the first one fails. Say the study code, the participant code and the date at the start, so the file identifies itself even after someone renames it. Our own Speech-To-Text shows why the room matters: clean studio audio transcribes well, and phone recordings with background noise are noticeably less accurate.
Get a first draft
You may already own a tool that does this. Word transcribes 300 minutes of uploaded audio a month for Microsoft 365 subscribers, according to Microsoft’s support page in 2026, and labels the voices “Speaker 1” and “Speaker 2”. Zoom transcribes cloud recordings on paid accounts and writes “Unknown Speaker” where it cannot tell who is talking. Teams transcripts carry timestamps and speaker attribution and live in OneDrive and SharePoint. Google Meet transcribes eight languages on selected paid Workspace plans, according to Google’s help centre in 2026, and saves to the organiser’s Drive. All four save you the typing. All four also keep the recording in the vendor’s cloud. And none of them writes your participant codes: you get real names, or a generic label when the software cannot tell. The general method is in how to transcribe video to text.
Fix speakers and timestamps
Automatic speaker separation saves you marking every turn by hand. It also makes mistakes. The model card of pyannote 3.1, an open reference model, lists diarization error rates from 7.8 to 50.0 percent across nine benchmark sets, scored without a forgiveness collar. The spread follows the recordings: the harder the audio, the higher the rate. So listen to every speaker change. The typical errors are two short turns merged into one, and one long turn split between two voices. Group interviews need this pass most, because every extra voice adds a boundary the model can miss. Fix the labels, then set timestamps at every turn or at your fixed interval. Machine timestamps mark segments, and a segment is not always a turn. Check that each time sits where the new speaker starts.
Style pass, anonymise, check, store
The last four steps are quick to describe and slow to do. Apply the style you chose, the same way in every file. Replace names, places and employers with codes, and log each replacement in a separate file. Then listen again against the text, ideally with a second person. Read along at normal speed and stop only where text and audio disagree. Mark what you still cannot hear as [word?] instead of guessing. Store the result with limited access in a known location. If the interview will be published, send it to the interviewee first whenever possible, as the Oral History Association’s 2018 principles ask.
In short: draft fast, then correct slowly. The time goes into speakers, style and the second listen.
Recorded your interviews on video or audio? Upload them to alugha, let our Speech-To-Text separate the voices, then relabel them with the template above. Create your account
How long an interview transcript takes
Start with the words.
By hand
Across 2,438 recorded telephone conversations of the Switchboard corpus, Yuan, Liberman and Cieri (2006) measured 196 words per minute over the whole conversation. Added up turn by turn, the rate was 164. An hour of interview therefore holds roughly 9,800 to 11,800 words, by arithmetic rather than by count. The average typist manages 52 words per minute, according to a 2018 study of 168,000 volunteers (Dhakal and colleagues). Speech outruns typing three to four times over, before anyone rewinds.
Dresing and Pehl’s handbook (2024) turns this into a planning figure: five to ten times the recording length for a simple transcript, proofreading included. Their fastest measured speed was 1:3. Linguistic notation takes far longer, about 18 hours per hour for a basic GAT2 transcript. Britten’s 1995 paper on qualitative interviews in the BMJ gave six to seven hours per taped hour. Rules of thumb such as “four hours per hour” circulate without a source, and Indeed’s seven to ten hours for beginners names none either. The handbook and Britten show where their numbers come from.

With an AI draft
An AI draft of clean audio is fast, and often good. In OpenAI’s Whisper paper (Radford and colleagues, 2022), the large-v2 model reached a word error rate of 2.7 percent on clean read speech. On two sets of telephone conversations it made 13.8 and 17.6 percent, and on a single distant meeting microphone 36.4. Word error rate counts substituted, deleted and inserted words against the words of a human reference transcript, as NIST’s scoring tool does. A claim of 98 percent accuracy therefore means about two wrong words in a hundred. In a one-hour interview, that is roughly 200 to 235 corrections.
Errors of another kind exist too. Koenecke and colleagues (FAccT 2024) found invented phrases in about 1 percent of Whisper API transcriptions, across 13,140 segments from 437 speakers. The saving is still real. One local, offline workflow cut transcription time by up to 76.4 percent across 12 interviews (arXiv, 2025). One workflow, not a general rate.
Vendor figures have one use: they tell you what a service is prepared to stand behind. They are promises, not measurements. Rev’s page promised in 2026 that its human service is 99 percent accurate or better. Otter’s product page said in 2026 that its AI generally reaches 85 to 90 percent on clear audio. Neither tells you what your café recording will produce.
A study budget
Now multiply. Hennink and Kaiser’s 2022 review of 23 studies found saturation at 9 to 17 interviews. At five to ten hours per one-hour interview, a study of that size needs roughly 45 to 170 hours of transcription by hand. That is arithmetic, not a measurement. Guest, Bunce and Johnson reported saturation within 12 interviews in 2006, from one sample of 60 interviews, so treat the low end as a best case. The second listen comes on top. It runs at least as long as the recording, so add one more hour per interview hour for every person who checks. Budget only the typing, and you budget too little.
In short: budget five to ten hours per interview hour by hand. An AI draft moves the hours into checking; it does not remove them.
Consent, anonymisation and GDPR
An interview recording identifies a person. So does a transcript with names in it. The GDPR defines personal data in Art. 4(1) as any information relating to an identified or identifiable natural person, and a recording of a voice is, in practice, personal data.

Consent before the first question
Record the consent itself, at the start of the recording. One exchange is enough: the interviewer asks whether the participant agrees to recording and transcription, and the participant answers in their own words. In Germany, § 201 StGB makes recording someone’s non-public spoken word without authorisation punishable by up to three years in prison or a fine. California’s Penal Code § 632 requires the consent of all parties to record a confidential communication. Other jurisdictions differ. Interviews about health, faith or politics touch the special categories of Art. 9 GDPR. Processing them needs explicit consent, unless the research exemption of Art. 9(2)(j), backed by national law such as § 27 BDSG, applies. For research, Recital 33 allows consent to certain areas of scientific research when the exact purpose is not yet clear.
Pseudonymise, log, do not black out
Replace. Do not black out. The UK Data Service’s 2023 guidance separates direct identifiers, such as name, address, telephone number and voice, from indirect ones, such as age, occupation and region. It asks for pseudonyms or replacements instead of blanks, an anonymisation log of every change stored separately, and no over-anonymising. A transcript that says P07, with a key file elsewhere, is pseudonymised under Art. 4(5) GDPR, not anonymous. In Germany, § 27 BDSG adds that special-category research data must be anonymised as soon as the research purpose allows, with identifying features stored separately until then. The voice itself counts as a direct identifier. If a participant must not be recognisable at all, the recording is the problem, not the transcript.
Where the recording goes when you upload it
US tools are fast, and many of them are lawful choices. Since 10 July 2023, a US vendor certified under the EU-US Data Privacy Framework can receive personal data from the EU. Otter’s privacy policy, for instance, says it stores data on Amazon Web Services in the United States, holds that certification and trains its models on de-identified audio and transcripts. For a public interview series, YouTube’s reach is unmatched as well. A confidential interview raises a different question: where the audio sits, and who can reach it. Any transcription or hosting service that processes the recording on your behalf is a processor under Art. 28 GDPR and needs a contract with you. The hosting side is covered in GDPR-compliant video hosting.
In short: consent on the record, pseudonyms with a separate key, and a processor contract for every service that touches the audio.
Multilingual interviews and dissertations
Two situations change the file layout.
Two languages, two files
Transcribe in the language that was spoken. Translate into a second file, never over the first. Code on the original, and quote the translation with the original in a footnote or appendix. Keep participant codes and timestamps identical across both files. A co-author can then jump from a line in the translation to the same second of audio. Translate the clean version, not the full verbatim one, because fillers and false starts do not carry across languages. Where a phrase only makes sense in the original, keep it and add a literal gloss in square brackets. Name the files so the pair stays together, such as RTW-2026_P03_DE and RTW-2026_P03_EN. Translation is one place where we help, within limits the next section sets out.
Transcripts in a dissertation
A dissertation rarely prints every transcript in the body. Full transcripts usually go into an appendix, or into a data archive agreed with your ethics board. Quote with the participant code and a paragraph number, such as P03, para. 14. Quotations from your own participants are part of your data, identified by their code, not entries in your reference list; check your department’s style guide for the exact form. Keep audio and transcripts for the retention period your funder sets. Guideline 17 of the DFG’s code of conduct, in force since 1 August 2019, says the research data behind a publication are generally archived for ten years. That is the German rule; other funders set their own. Your ethics approval may also limit whether full transcripts can appear in print at all, so check it before you plan the appendix.
In short: one language per file, codes and times identical across both, and quotations cited by code and paragraph.
Interview transcripts with alugha
We are one answer here, not the default one. Start with what we do not do.
What we do not do, and when you do not need us
Our output is time-coded segments, each with a voice, not a formatted research transcript. The header, the labels, the tags and the style pass are your work, and the template above is for exactly that.
Our Speech-To-Text separates the voices. Naming them is still your job.
There is no anonymisation feature: you replace names by editing segments in the dubbr, our editor, and the audio stays untouched. There are no verbatim modes either. Transcription uses credits, and exporting the file sits on a paid tier. The Plain Text export carries no timestamps, and there is no bulk export: one language and one format per download. We have no integration with NVivo, MAXQDA or ATLAS.ti; you export the text and bring it in yourself.
Some interviews do better without us. A conversation-analysis transcript typed by hand, where every pause is data. A single audio interview that never leaves your laptop, which a local, offline workflow handles well. A recording that your ethics approval says no processor may receive. Stay offline and use the template.
Where we fit
We fit best with a series of interviews, audio or video, two or more speakers, and a team that needs recordings, transcripts and access rights in one place. Our upload wizard takes nearly any audio or video file. Switch on Process audio only for an audio interview, and the encoding step costs no credits. We do not cap the length of a single upload, so a ninety-minute interview goes in as one file.
Our Speech-To-Text then transcribes the language track you select, separates the voices and gives each one its own colour. The segments land in the dubbr, where you edit them one by one: a misheard name, a sentence split in two, a real name replaced with a code. The steps are in transcribe a video with Speech-To-Text.

Export the transcript as Plain Text, WebVTT or SRT; the last two keep every segment’s timing. The options are in export subtitles.

A project can stay Private while you work on it. The three states are Private, Public and Not listed. On a paid tier, folder-level permissions decide which colleagues reach which set of interviews. Recordings and transcripts sit on our hosting in Germany and the EU, with no third-party advertising cookies and trackers. Once the transcript is right, we translate it into the languages your co-authors read, with a glossary on a paid tier for names and terms.

Publishing a video interview on your own site adds one rule: under § 25 TDDDG, anything a page stores or reads on a visitor’s device needs consent unless it is strictly necessary. The embed matters as much as the host.
In short: we separate voices and keep recordings and transcripts in one place, hosted in Germany and the EU. The formatting stays with your template.
Frequently asked questions about interview transcription examples
What does an interview transcript look like?
An interview transcript opens with a header naming the study, participant code, date, style and timestamp rule. Every turn then starts with a timestamp and a speaker label, such as [00:03:20] P03, followed by the words and bracketed tags for what they miss. The example at the top shows both styles; the video transcript examples cover subtitle formats.
How long does it take to transcribe a one-hour interview?
By hand, plan five to ten hours per hour of audio, proofreading included, following Dresing and Pehl’s 2024 handbook. Britten’s 1995 BMJ paper gave six to seven. An AI draft shortens the typing, not the checking: one local workflow saved up to 76.4 percent of the time across 12 interviews in 2025. Budget the range, not the best case.
What is the difference between verbatim and clean verbatim?
Full verbatim writes down every word and sound: fillers, stutters, false starts, timed pauses and the interviewer’s mm-hm. Clean verbatim removes those and keeps every word of substance in the speaker’s order. Neither style paraphrases. Use full verbatim when how something was said is the data, as in conversation analysis, and clean verbatim when what was said is the data.
How do I cite interview transcripts in a paper?
Cite a quotation by participant code and paragraph number, for example P03, para. 14, and put full transcripts in an appendix or an agreed data archive. Timestamps are for finding the audio, not for citing it. Quotations from your own participants are part of your data, not entries in your reference list. Your department’s style guide sets the exact form.
Should a transcript include ums, laughter and pauses?
It depends on the style. Full verbatim keeps every um and times each pause in seconds, such as (2.0). Conversation analysts write (.) for a micropause under two tenths of a second. Clean verbatim drops fillers and pauses. Laughter stays in either style when it changes the meaning. Write the rule you follow, and every symbol, into the header.
Should I use real names in an interview transcript?
Not in research. Use codes such as P07 and keep the key linking codes to names in a separate, restricted file; the transcript is then pseudonymised, not anonymous. HR records may keep names under strict access control. A published interview can name the speaker, with consent and after the speaker has reviewed the text.
How do I mark inaudible passages and overlapping speech?
Write [inaudible] with the time, such as [inaudible 00:04:12], so anyone can find the gap in the audio. Mark a guess as a guess, for example [Kiel?]. For crosstalk, a simple [overlap] tag at both turns works for most studies. Conversation analysis is stricter: a left square bracket marks where an overlap starts, a right one where it ends.
Can I use AI transcription and clean it up myself?
Yes, and it is the common workflow now. Check every speaker change, because automatic speaker separation makes mistakes, and check names and technical terms against the audio. AI models also occasionally invent phrases nobody said. Before you upload a recording, make sure the service is bound by a processor contract under Art. 28 GDPR and that your consent form covers it.
Getting started
Go back to the sentence from the start. Here it is in the finished template.
STUDY RTW-2026, return to work after long illness PARTICIPANT P03 STYLE clean verbatim TIMESTAMPS at every speaker turn REPLACEMENTS logged in RTW-2026_key.xlsx, stored separately [00:03:20] P03: I didn't want to go back.
The hesitation is gone. The decision is not, because the header says what was removed, and the timestamp says where the audio is. Anyone on the team can hold that line against the recording and see the same thing you saw. Pick your style, copy the template, transcribe the first interview, and do the rest the same way.
Running an interview series with a team? Talk to us about keeping recordings and transcripts in one place, hosted in Germany and the EU, or create your account and upload the first one today.



