A global training team once recorded twelve product demo videos in English, sent the raw files straight to a translation vendor, and got a quote back that tacked on three unplanned days — for transcription nobody had budgeted, because nobody realized it was a separate step at all.
The Step Everyone Forgets to Budget For
It’s an easy mistake to make, because the missing step is invisible until it isn’t. A recording feels finished the moment it’s recorded. But translators work from text, not audio — so somewhere between “we have the video” and “we have the translated version,” someone has to turn spoken words into a clean written transcript.
If all you have is a recording, someone turns it into words first — either you, unexpectedly, or your vendor, at an unplanned cost.
Transcription vs. Captioning vs. ASR: Three Things That Get Confused
Transcription, captioning, and automatic speech recognition often get lumped together. They’re related, but they solve different problems. The American Translators Association treats a clean source transcript as the starting point for any downstream language work — and the distinctions below explain why:
Where This Shows Up Most Often
A few situations where transcription quietly becomes the bottleneck if it isn’t planned for upfront:
- Multilingual subtitles — every translated caption track starts with an accurate source-language transcript.
- Webinars and training libraries going global — one clean transcript can become subtitles, a dubbing script, and a written guide, all from a single source.
- Legal depositions and recorded interviews — often need a certified transcript before any translated version is usable at all.
- Research interviews and focus groups — teams need transcripts to code responses; translated transcripts let global research teams work from one dataset.
What Separates a Usable Transcript From a Rough One
Three things are worth checking before you commit to a provider — everything else is negotiable, these usually aren’t: human review on top of any automated pass, speaker labeling and timestamps for anything with more than one voice, and confidentiality handling as standard practice for legal, medical, or internal recordings — not something you have to ask for separately.
The teams that handle this most smoothly treat transcription as the first step in the pipeline — they request it the moment they record content, not weeks later when a translation request suddenly surfaces the need.
Not sure if your project needs a transcript first?
It’s usually a quick answer — send us the file and we’ll tell you straight.
Ask Our Team
