Speech to Subtitles

Generate SRT/VTT subtitles from the audio inside your video — Drop in a video (MP4, WebM, MOV, MKV, AVI, …) or a plain audio file, and a local neural speech engine (Chinese + English) transcribes its audio track directly with accurate timestamps. Everything runs in your browser — no microphone, no upload, no cloud. First use downloads the ~95 MB neural engine once, then it is cached.
Drop a video file here, or click to choose
MP4 · WebM · MOV · MKV · AVI · MPG · FLV · MP3 · WAV · M4A — up to 200 MB
Local engine: not loaded.
Choose a video file to begin.
Subtitles (editable — adjust the text, then download)
No subtitles yet.
What is automatic subtitling?

Automatic subtitling runs speech recognition over a video’s audio track and turns the recognized words into timed cues: each stretch of speech becomes a short subtitle block with a start and end timestamp, exported as SRT or VTT. The hard part is segmentation — a transcript is a wall of words, but viewers can only read a few words at a time, so long utterances must be split at natural pauses into cues that stay on screen long enough to read.

SRT and VTT are both plain-text cue lists that differ mainly in timestamp style and header — players that accept one almost always accept the other, so download whichever your player asks for.