Speech is the input
Transcription listens to the audio track. It does not extract titles, signs or existing burned-in subtitles from the picture. A silent recording cannot produce a genuine spoken transcript.
Get words from the speech in an uploaded recording. Review the transcript against playback, correct it, and download TXT, SRT or VTT.
Up to 100 MB · 5 minutes · files expire after 24 hours
English speech recognition. Review names and timing against playback.
Review the words and timing before you export. Timed text and burned-in video are different outputs.
Saved local English draft · October 4, 2026
Welcome to Frame and Voice. Today we will resize a landscape video without stretching.
Keep the circle round, place captions above the bottom controls, and listen before you export.
Actual local speech result, October 4, 2026: faster-whisper-tiny transcribed the owned clip into two timed segments. The opening reads “Welcome to Frame and Voice. Today we will resize a landscape video without stretching.” The two-second silence fixture produced no cues or transcript. These fixtures do not establish general accuracy.
Upload a recording with audible speech.
Run speech recognition and review timed segments.
Correct errors and download your chosen text format.
Transcription listens to the audio track. It does not extract titles, signs or existing burned-in subtitles from the picture. A silent recording cannot produce a genuine spoken transcript.
Names, accents, noise and overlapping speech can cause errors or invented text. English speech is the tested launch scope. Speaker identification is not promised. TXT suits reading; SRT/VTT retain timing for captions. Listen and correct the output before use.
Sources & context
Source pages retrieved 2026-10-04. Examples labelled illustrative are explanations, not saved processing results.
Import or create timed captions, edit their appearance, and export a subtitled MP4.
Convert SRT and WebVTT, inspect cue warnings, and download a timed subtitle file.
Reduce steady noise and normalize toward −16 LUFS, compare aligned playback, and download real audio.