Upload your file
Choose a supported audio or video recording up to 100 MB.
Upload your file for free Interview Transcription. Review synchronized words, detected speakers and sound events, edit the text and export the result.
Long recordings can take a few minutes. Keep this tab open.
Review an interview with timestamps instead of repeatedly scrubbing the recording. The editable transcript helps journalists, researchers and recruiters locate quotes quickly.
Choose a supported audio or video recording up to 100 MB.
The system detects language, speakers, words and sound events, then aligns them with timestamps.
Play it back, edit the text and download the format you need.
Read along with the recording, copy clean text, or export files for documents, captions and your own workflow.
Click any word to jump to its exact moment and follow the active phrase during playback.
Different speakers are identified automatically, while events such as laughter or applause can be marked when detected.
Automatic mode can recognize speech across 90+ languages; seven common languages also have manual presets.
00:12 — “Click any word to return to the matching moment in the recording.”
Raw uploaded media is deleted after processing. Transcript text and technical request data may be retained in protected operational logs for service reliability, abuse prevention and support.
Privacy and retention →Yes. The converter and all listed export formats are available without creating an account.
Automatic mode recognizes more than 90 languages. Speaker information and events such as laughter or applause are added when detected.
Microphone quality, background noise, compression, accents and overlapping speakers can all affect the result.
Yes. You can transcribe and download results without a subscription or account.
MP3, WAV, M4A, MP4, FLAC, AAC, OGG, OGA, OPUS, WebM, WMA, AVI, MKV, MOV, MPEG and MPGA are accepted up to 100 MB.
Automatic detection covers more than 90 languages. English, Spanish, Portuguese, German, French, Russian and Turkish can also be selected manually.
Speaker information is added automatically. Sound events such as laughter or applause may also appear when detected; structured details are included in JSON.
Yes. Every result includes both SRT and VTT downloads with timestamps.
Accuracy depends on recording quality, background noise, compression, accents and overlapping speech. Verify important names, numbers and quotations against the recording.