Audio and video transcription Transcription services: from audio to text.
Interviews, conference recordings, podcasts, training videos: our audio and video transcription combines AI with a check by native-speaker editors. Add specialist translation into 220+ languages if you need it.
Over 140,000 customers trust tolingo

- Service: Turning audio and video files (interviews, conferences, podcasts, training) into text, with AI-assisted transcription and human Medial Post-Editing. Add specialist translation if you need it.
- Quality: Native-speaker editors check every transcript for accuracy, timestamps and speaker attribution. ISO-certified processes.
- Cost: AI from €1.50/minute, Smart (AI + post-editing) from €6.50/minute, Premium (+ translation) on request.
- Turnaround: AI: 1 working day. Smart: 3 working days. Express available.
What is a transcription?
Definition and common use cases
A transcription turns spoken content into text. In practice, that means an audio or video file becomes a formatted document with timestamps, speaker labels and, where needed, notes on non-verbal events. Read more in the Transcription glossary entry.
Common use cases: interviews for market research, court records, conference recordings, podcast episodes for accessibility and SEO, training videos and academic research interviews. In every case, a recording becomes a searchable, citable text.
Three packages, one outcome: your finished transcript
AI, Smart or Premium, depending on what you need
We offer three service levels. All packages accept common audio and video formats (MP3, WAV, FLAC, M4A, MP4, MOV, AVI, WebM and more). You receive the transcript as a Word file (.docx), and on request also as .srt (subtitles), .txt or .pdf. Full pricing information is also available on our prices page.
Automatic AI transcription. Best for clear recordings with few speakers and standard terminology.
Request AI transcriptionAI plus human Medial Post-Editing. Best for technical terms, dialects, multiple speakers.
Request Smart transcriptionSmart plus specialist translation into one or more target languages. From one interview to 220+ languages.
Request the Premium packageFor everything at a glance, download the PDF: Transcription one-pager (PDF).
What is a transcription used for?
Typical projects from our customers
Focus groups, in-depth interviews, user surveys. Transcripts form the basis for qualitative content analysis.
Witness statements, hearing transcripts, due diligence interviews. Citable, and under NDA if you need it. More on legal translations
Research interviews, expert conversations, ethics board recordings. Transcripts for qualitative studies and dissertations. More on academic translations
Podcast episodes, video interviews, press conferences. Transcripts for accessibility, SEO and content repurposing.
Ready to commission a transcription?
Send us your audio or video file. We'll recommend the right package and give you a price and delivery date.
Free · No obligation · NDA on requestWhy AI alone isn't enough
What human post-editing adds
Modern speech recognition can transcribe audio in minutes. On a clear recording with one speaker, accuracy runs at 90-95%. But technical terms, dialects, background noise, speaker changes and proper names remain difficult. In a medical interview, the AI might mishear "baseline" as "base light." In a legal context, "tort" can become "torte." That's why a well-maintained termbase pays off on transcription projects too.
That's why, in the Smart and Premium packages, native-speaker editors check every transcript during Medial Post-Editing (MPE). They correct technical terms, add timestamps, attribute speakers and apply the tolingo rule set (basic or extended transcription rules). The result: a citable, searchable text.
Basic and extended transcription rules
Which rule set fits your project
tolingo works to two standardised rule sets. You choose which one applies when you place your order.
Basic transcription
The text is captured exactly as spoken, including errors and filler words. Unclear passages are marked [inaudible], other languages as [foreign language]. Speakers are labelled "Speaker 1", "Speaker 2" and so on. No timestamps.
Extended transcription
Everything in basic transcription, plus: dialect colouring is corrected, non-verbal sounds (ums, stutters) are removed, listener acknowledgements are kept only where they add content. Notable events go in brackets (e.g. "(audio disruption)"). Sensible punctuation for long sentences. A timestamp before every speaker change. Pauses over 4 seconds marked with their length. Unclear words flagged with a question mark: (?landmark).
For a detailed practical guide with examples, see our guide: Avoiding mistakes in interview transcription.
Transcription and translation, one team
From the spoken word into any target language
In the Premium package, we deliver transcription and specialist translation as a single project. That saves you coordinating between a transcription provider and a separate translation agency: one project manager handles both.
A typical scenario: an international company has training videos transcribed in German, then translated into English, French and Spanish. The transcript becomes the source text directly for translation into 220+ languages. On request, we also produce subtitle files (.srt) for your video platform.
Ready for your transcription?
Interview, conference or podcast: we transcribe and translate your content.
Free · Reply within hours · 97.8% customer satisfactionGlossary
Terms around transcription and speech processing
- Medial Post-Editing (MPE)
- Human review of an AI-generated transcript. The editor checks technical terms, corrects speaker attribution, adds timestamps and applies the agreed rule set. Comparable to post-editing in machine translation.
- Transcription
- Turning spoken content from an audio or video file into text. The result is a formatted document with speaker labels, optionally with timestamps and notes on background noise or pauses.
- Timestamp
- A marker in the transcript showing exactly when a statement occurs in the original recording (e.g. [05:33]). Makes it easy to find passages in long recordings, and standard on extended transcriptions.
- SRT file (SubRip Subtitle)
- The standard subtitle format. Contains numbered text with start and end times. Supported by YouTube, Vimeo and most video platforms. tolingo can also deliver transcripts in .srt format on request.
- Speaker diarisation
- Automatic or manual attribution of passages to individual speakers in a recording. In AI transcripts, shown as "Speaker 1", "Speaker 2" and so on. During post-editing, these are replaced with the speakers' actual names where known.
Frequently asked questions about transcription services
Answers to the questions we hear most
What does a transcription cost at tolingo?
AI transcription from €1.50/minute, Smart (AI + post-editing) from €6.50/minute. The Premium package (transcription + translation) is available on request. The exact price can vary by language and audio quality. Every quote comes with a fixed price.
How quickly will my transcription be ready?
The AI transcription itself takes only a few minutes. With project management and post-editing added, expect 1 working day (AI), 3 working days (Smart) or 4+ working days (Premium). For urgent projects, we offer Express and Super Express.
Which audio and video formats do you accept?
Audio: MP3, WAV, FLAC, M4A, AAC, DSS, GSM, MP2, MPA, OGG. Video: MP4, MOV, AVI, WMV, MXF, FLV, MPEG, MPG, 3GP, WebM, M4V, TS. You receive the transcript as a Word file (.docx), and on request also as .srt, .txt or .pdf.
Which languages do you offer transcription in?
Over 25 languages, including German, English, French, Spanish, Italian, Portuguese, Dutch, Polish, Russian, Turkish, Chinese, Japanese, Korean, Arabic and all Scandinavian languages. Ask us for the full list, or see our languages page.
Can I have the transcription translated too?
Yes. Our Premium package combines transcription and specialist translation in one project. The transcript becomes the source text directly for translation into one or more target languages (220+ language pairs).
How confidentially do you handle my recordings?
All files are transferred over SSL-encrypted connections. Our information security management is ISO 27001 certified. On request, we sign a separate non-disclosure agreement (NDA).
Basic or extended transcription: what's the difference?
Basic transcription captures the spoken text word for word, including filler words and grammatical slips. Extended transcription corrects dialect colouring, removes non-verbal sounds, adds timestamps and marks pauses. See our interview transcription guide for details.
How do I order a transcription?
The easiest way is by email or contact form. Send us your file (or a short sample for a test quote), and tell us the language and the rule set you want. We'll come back with a fixed-price quote and a delivery date.
