Sonomir

Guides · 5 min read

How to transcribe an interview without spending your evening on it

Typing an interview by hand takes four to six hours per recorded hour. Here is the workflow that gets the same result in about forty minutes, and the three places where you still have to do the work yourself.

The short version

  1. Record at arm's length, one file, no fancy setup.
  2. Run the file through automatic transcription and get a timestamped draft.
  3. Read the draft while skimming the audio at 1.5x, fixing names and numbers only.
  4. Pull quotes with their timestamps into a separate document.
  5. Keep the audio. It is your only defence if a quote is disputed.

Steps 1, 3 and 5 are yours. Step 2 is the part that used to cost four hours and now costs minutes, and step 4 is where the actual thinking happens.

Recording: the only decision that changes the outcome

Transcription accuracy is set by microphone distance more than by anything else. A phone lying between two people on a table is roughly a metre from each mouth; a phone held in the hand of whoever is speaking is twenty centimetres away. That difference is worth more than any model choice, any noise-reduction setting and any amount of post-processing.

For an in-person interview, put the phone on the table in front of the person you are interviewing, not halfway between you. Your own questions will be a little quieter and you already know what you asked.

For a remote interview, record locally rather than relying on the platform's cloud recording. Zoom, Teams and Meet all compress aggressively for bandwidth; a local recording keeps more of the voice, and a separate audio track per speaker — which Zoom can be told to produce — makes speaker labels almost perfect.

Two things not to bother with: an external microphone you have not used before, and a "voice enhancement" setting. Unfamiliar gear fails in ways you discover afterwards, and aggressive noise suppression removes consonants along with the hum.

Transcription: what to expect, honestly

A clear recording with one or two speakers comes back at roughly 92–97% word accuracy. That means one wrong word every couple of lines, and the errors are not random: they cluster on proper nouns, technical terms, numbers and anything said while someone else was talking.

Turn on speaker labels for any recording with two or more voices. An interview transcribed as one undivided block is barely usable — you cannot tell a question from an answer without listening again, which defeats the point.

Expect a 40-minute interview to take a minute or two to process. If a tool promises "instant" on an hour of audio, it is either queueing the work and showing you a spinner or transcribing at a quality you would not accept.

You can do this on the transcription tool here: ten minutes a day are free with timestamps and a summary, and a longer file is priced per minute. If your interview is already on YouTube, the transcript tool reads the caption track directly instead.

The clean-up pass, and how to keep it to twenty minutes

Do not fix the transcript. Fix the parts you are going to use.

Open the transcript and the audio side by side, play at 1.4–1.6x, and read along. You are looking for four things and ignoring everything else: names, numbers, technical terms, and any sentence you might quote. Those are the four categories where an error changes meaning rather than just looking untidy.

Leave the filler. Every "um", every false start, every repeated word — leave them until you are actually writing, then clean up only the quotes you use. Cleaning a full transcript you will use 5% of is the classic way to turn a forty-minute job back into a four-hour one.

Keep the timestamps in the draft. They are what lets you jump back to the moment when an editor asks whether someone really said that. Strip them from the final document, not from your working copy.

Quotes: pull them with context, not alone

Copy each usable quote into a separate document with three things attached: the timestamp, the question that produced it, and the sentence that came after it.

The question matters because a quote's meaning often lives in what was asked. The following sentence matters because that is where people qualify what they just said, and a quote that drops the qualification is the most common way an accurate transcript still produces an inaccurate article.

Mark ellipses honestly. If you cut words from the middle of a quote, the reader should see that you did.

If your interview is going to become teaching material rather than an article — a lecture summary, a revision sheet — the transcript is also the cheapest possible source of questions. That is the one thing worth automating after the transcript exists.

What automatic transcription still cannot do

It cannot certify. Court filings, clinical records and anything with a legal chain of custody need a human transcriber and a certificate of accuracy; Rev and similar services charge around $2 a minute for exactly that, and they are the right answer when you need it.

It cannot fix a bad recording. If two people talked over each other for ten minutes, no tool separates them reliably, and the honest output is a mess with speaker labels drifting.

It cannot tell you what mattered. A summary gives you the shape of a conversation; the decision about which forty seconds of ninety minutes to publish is the job.

Everything else — the four hours of typing — is solved, and has been for about two years. If you are still typing interviews by hand, that is the habit to drop this week.

Questions

How long does it take to transcribe a one-hour interview?
Automatically, a couple of minutes of processing. Then budget 20–30 minutes for the pass that fixes names, numbers and quotes. By hand it is four to six hours, which is the comparison that matters.
Is automatic transcription accurate enough to quote from?
Not without checking the specific sentence you are quoting. It is accurate enough to find the quote, and you verify that one sentence against the audio before it goes anywhere.
Should I record the interview on Zoom or on my phone?
Both, if it is remote: the platform recording as a backup, and a local recording as your working file. If you can enable a separate audio track per speaker, do — it makes speaker labels near-perfect.
Do I need to transcribe the whole interview?
You need a full timestamped draft to find things in, which automatic transcription gives you cheaply. You do not need to clean the whole thing — only the parts you publish.

The machines this uses

Related guides