Drop in a video or audio file and get a downloadable SRT or VTT, transcribed by Whisper running on your own machine. Your media is never uploaded.
This tool runs a WebAssembly engine and must be opened over a local server, not straight from a file.
In this folder run python3 -m http.server 8777, then visit http://localhost:8777/transcribe.html.
One-time setup needs internet: the first run downloads the speech model (≈40–80 MB) and the transcriber library, then caches them in your browser. Cached downloads may be reused, but offline availability is not guaranteed. Your audio itself is processed locally and never leaves your device.
Drop a video or audio file, or click to choose
MP4, MOV, WebM, MP3, WAV, M4A… — stays on your device