What is transcription: the term defined, with an honest account of how accurate it is
Transcription is the conversion of recorded speech into written text, stored alongside the recording so that the spoken part of a submission can be read, searched and quoted instead of replayed in full. In workflow software it is normally produced automatically by speech recognition, also called speech to text or ASR.
The point of it in an operations context is not the text itself. It is that a spoken explanation stops being a five minute video nobody has time to open and becomes something a colleague can scan in twenty seconds and search across a year of cases.
What does transcription mean here?
Two related outputs both get called transcription, and they behave very differently.
- Automatic transcription. Machine generated, available within seconds or minutes, usually timestamped so a phrase links back to its point in the audio. Accuracy varies enormously with conditions.
- Verified transcription. Machine output corrected by a person against the recording, or produced by a trained transcriber from the start. Slower, more expensive, and the only kind that can carry evidential weight.
Diarisation, the separation of who said what, is a distinct capability from transcription and frequently the weaker one whenever several people are speaking.
Ready to see what your customers see?
Send one link. Get guided, verified video back. No app, no account.
How does automatic transcription work?
The audio is split into short frames and turned into acoustic features. A model maps those to the most probable sequence of words, weighing what the sound resembles against what is plausible in the language. The output arrives as text with time offsets and, in most systems, a confidence value per word or per segment.
Confidence is worth showing to reviewers and is routinely hidden. A word the system flagged as uncertain, displayed as uncertain, gives a reader something they can use safely. The identical word rendered in plain text is a fact the reader has no reason to doubt.
How accurate is transcription?
Accuracy is reported as word error rate: the proportion of words inserted, deleted or substituted when measured against a reference transcript. On clean, single speaker, scripted audio, modern systems are strong. Those are not the conditions in which a customer records anything.
- Accent and dialect. Koenecke and colleagues, publishing in PNAS in 2020, tested five commercial systems from Amazon, Apple, Google, IBM and Microsoft on matched interviews and found an average word error rate of 0.35 for Black speakers against 0.19 for white speakers. Roughly double the errors, on the same task, from the same systems.
- Noise and multiple speakers. The CHiME 5 challenge used real dinner party recordings, with overlapping speech and household noise. The published baseline scored 81.1 percent word error rate, and a Johns Hopkins system built specifically for the task improved that to 69.4 percent on the development set. Under those conditions most of the words are wrong.
- Technical vocabulary. Part numbers, model codes, chemical names and local trade terms are simultaneously the words a reviewer most needs and the ones least likely to be well represented in training data.
- Recording conditions. A phone held at arm's length in a workshop, outdoors in wind, or next to a running engine is the normal case in this line of work, not the difficult one.
For comparison, the human professional standard: the National Court Reporters Association requires 96 percent accuracy at 200 words per minute to certify a realtime reporter, on clean two voice question and answer material.
When a transcript is not a record
An automatic transcript is a working aid. It is not a verbatim record, and it should not be handled as one unless a person has checked it against the audio and is prepared to say so.
That matters most in exactly the places transcripts are most useful. A claims handler quoting a customer's account of how the damage happened, or a warranty assessor citing what a technician said about a fault, is leaning on words a machine guessed at. The recording is the record. The transcript is an index into it.
- Keep the audio. A transcript that cannot be checked against its source can never be corrected. Retention policy has to cover the recording, not only the text.
- Quote from the audio, not the transcript. If a phrase is going into a decision letter or onto a file, listen to it first.
- Mark corrections as corrections. An edited transcript should show that it was edited, by whom and when, the same as any other case record. That is the standard set out under evidence integrity.
- Never let the transcript decide. Search and triage on the text if it speeds things up. The assessment stays with the qualified person, who watched the recording.
Where transcription earns its place
The real gain is search. One transcribed submission saves a few minutes. A year of them makes a spoken archive queryable, so a fault description that keeps recurring, or a phrase that keeps appearing shortly before a dispute, becomes findable at all. That is a case management capability more than a media one.
Speech to text sits alongside the video and photos in the case file in Venta Capture, a product of VentaVid, where a customer's spoken explanation is transcribed and searchable next to the recording it came from. The accuracy limits above apply to it exactly as they apply to any other system, which is why the recording, not the transcript, stays the record. Related reading: warranty claim inspection.
Ready to see what your customers see?
Send one link. Get guided, verified video back. No app, no account.