Brazilian and European Portuguese
The two varieties sound different enough that human learners struggle to move between them, and transcription handles both — Brazilian Portuguese somewhat better, because there is far more of it in training data.
European Portuguese reduces unstressed vowels heavily, which makes word boundaries harder and produces more errors on function words. Brazilian Portuguese is more open and transcribes closer to Spanish-level accuracy.
Spelling follows the post-reform orthography. If your institution requires pre-reform spelling, plan on a find-and-replace pass.
For recordings that mix Portuguese and Spanish — common in border regions and in Latin American business calls — leave the language on auto-detect instead of using this page, since a fixed language setting will force the wrong one onto half the recording.
What this tool does
Upload an audio or video file and the transcript appears on this page. Every paragraph carries the timestamp it came from, so you can find the moment in the original recording without scrubbing.
Speaker labels are available when the recording has more than one voice: each turn is tagged Speaker 1, Speaker 2 and so on, which is what makes interview and meeting transcripts usable rather than a single undivided block.
Files never become public. The upload is processed, the transcript is attached to your job, and nothing is indexed or shared.
Formats that work
Audio: MP3, M4A, WAV, AAC, FLAC, OGG, Opus, WebM, AMR and WMA.
Video: MP4, MOV, WebM and MKV — the audio track is read and the video itself is ignored, so a 2 GB screen recording transcribes the same as the audio ripped from it.
Files up to 500 MB upload directly from your browser to storage, which means a long recording is not limited by any request size cap. A 90-minute MP3 at a normal bitrate is around 80 MB; an hour of video is usually 300–700 MB.
If your file is larger than 500 MB, export audio only from your editor first. It transcribes identically and uploads in a fraction of the time.
How the free tier works
Ten minutes of audio a day, across up to three files, cost nothing. You see the transcript, the timestamps, the summary, and you can copy all of it.
When a file is longer than your remaining free minutes, the transcript stops at that point and says exactly where it stopped. Nothing is charged and nothing is hidden behind a signup form — you get a real result for the first ten minutes rather than a blurred teaser.
To transcribe the whole file, add credits. Credits are bought once, never expire, and work across every tool on this site. At 1 credit a minute, a 45-minute interview costs 45 credits, which is about $1.93 from the middle pack.
Speaker labels and why they matter
A two-person interview transcribed without speaker separation is nearly useless: you cannot tell a question from an answer without listening again.
With speaker labels on, each change of voice starts a new paragraph tagged with a speaker number. Rename them in your own document afterwards — the DOCX download puts speaker names in bold at the start of each turn precisely so a find-and-replace turns Speaker 1 into a real name in two keystrokes.
Diarisation is not magic. Two people talking over each other, a poor phone line, or someone joining halfway through are all cases where labels drift. Fix the three or four places it matters and leave the rest.
What to expect on accuracy
On a clear recording with one or two speakers, expect 92–97% word accuracy. That means roughly one wrong word per two lines, usually a name, a unit, or a technical term.
Things that reliably hurt accuracy: recording from across a room, air conditioning or traffic, several people speaking at once, strong accents in a language the model sees less of, and music underneath speech.
Things that help, and cost you nothing: record at arm's length rather than across a table, use the phone's own voice memo app rather than a video call recording where possible, and say proper nouns clearly the first time they come up.
Before you publish or submit anything, check names, numbers and any sentence you intend to quote. The timestamps make that a minute of work rather than a re-listen.
Subtitles from a video file
Upload the video, download the SRT, drop it on the timeline. The cues are split on natural sentence boundaries with millisecond timing, which is what Premiere, Final Cut, Resolve, CapCut and every web player expect.
VTT is the same content in the format HTML5 video wants, with speaker names as voice tags so a player can style each speaker differently.
Subtitle files always need a human pass — line lengths, a mis-heard brand name, a laugh transcribed as a word. Editing an SRT that is 95% right takes a few minutes. Typing one from scratch takes an hour per ten minutes of footage.
Example result
interview-2026-03-14.m4a — 38 minutes, two speakers, phone microphone
Transcript: 52 timestamped paragraphs, alternating Speaker 1 / Speaker 2. Summary: two sentences plus 7 substantive points. Downloads: SRT for the video cut, DOCX for the quote sheet.
Cost: 10 free minutes, then 28 credits for the remaining 28 minutes — about $1.20.
How this compares
Where another tool is the better choice, it says so.
Otter.ai
BETTER THERELive meeting capture, a bot that joins calls, shared workspaces and searchable history across a team.
BETTER HERENo subscription and no account. Upload the file you already have, read the transcript, pay per minute only for what you transcribe. Better for the person with one interview, not a recurring meeting habit.
Rev
BETTER THEREHuman transcription at 99% accuracy with a certificate, which automated tools cannot match for legal work.
BETTER HERERev's human service is $1.99 a minute and takes hours. This is about $0.05 a minute and takes seconds, which is the right trade for research notes, subtitles and first drafts.
Whisper on your own machine
BETTER THEREFree forever, fully private, and you can run it on a whole folder overnight.
BETTER HEREIt needs a capable machine, a Python environment and patience, and it does not label speakers out of the box. This is the same class of model with nothing to install.
Your questions, answered.
- Brazilian or European Portuguese?
- Both, with no setting to change. Brazilian Portuguese is slightly more accurate because there is more of it in training data.
- Which spelling convention is used?
- The post-reform orthography. Pre-reform spelling needs a find-and-replace pass afterwards.
- Is it really free?
- Ten minutes of audio a day across up to three files, including timestamps and the summary, with no account and no email. Credits cover anything beyond that plus the SRT, VTT and DOCX downloads.
- What is the largest file I can upload?
- 500 MB. The upload goes straight from your browser to storage rather than through a request body, so long recordings work. Above 500 MB, export an audio-only version first.
- How long does it take?
- Roughly 10 to 30 seconds for a 10-minute file once the upload finishes, and a few minutes for a two-hour recording. Long jobs keep running if you close the tab — the page picks the result back up when you return.
- Can it tell speakers apart?
- Yes, when speaker labels are switched on. Each change of voice becomes a new labelled paragraph. It works best on recordings where people do not talk over each other.
- Which languages work?
- About 50, with the strongest results in English, Spanish, Portuguese, French, German, Italian, Dutch, Hindi, Japanese and Mandarin. Leave the language on auto-detect unless the recording opens in a different language from the one it continues in.
- Do you keep my file?
- The upload is kept only as long as the job needs it and is then deleted. Transcripts stay attached to your job so you can come back to them, and are never rendered on a public page or used to train anything.
- Can I transcribe a phone call or a voicemail?
- Yes, if you have the file. AMR and M4A voice memos, call recordings and voicemail exports all transcribe, though phone-line audio is narrower band than a microphone so accuracy is a few points lower.
- Can I upload several files at once?
- One at a time today. Bulk upload is a credit-priced feature on the roadmap; if you need it now, run them back to back — credits are charged per minute either way.
- What if the transcript is wrong?
- Copy it out and fix the handful of errors, which is the workflow every professional transcription service also assumes. If a whole file comes back garbled — wrong language detected, or silence — that is a bug worth reporting, and charged credits are refunded.
- Is this good enough for legal or medical work?
- No. Certified transcription for court or clinical records requires a human transcriber and a certificate of accuracy. Use this for research, notes, subtitles and drafts.