Upload audio or video, get timestamped SRT / WebVTT / TXT subtitles powered by Whisper running locally. Edit inline, preview on-video, download instantly.
Your audio/video never leaves your device. Whisper tiny model loads once (~230MB) and caches locally. Shares cache with AI Speech to Text.
Runs entirely in your browser. Your data never leaves your device.
When you click Generate Subtitles, four things happen — all locally, all in your browser:
SRT (SubRip Text), WebVTT and plain TXT. SRT works with almost every video player.
Yes. Both tools use whisper-tiny and share the same cached model in IndexedDB.
Yes. MP4, WebM and MOV files are supported. The audio track is extracted in your browser.
Yes. Each line shows start/end time and text. Click to edit inline before downloading.
We tested 30 mixed video/audio clips — average generation time 9.2s for 60s audio on MacBook Air M2.
| Metric | Value |
|---|---|
| Model size (INT8) | ~230 MB |
| Average processing speed | ~6.5x real-time (M2) |
| Model load (cold start) | ~9s (CDN) / <0.5s (cached) |
| Timestamp accuracy | ±0.5s |
| Formats output | SRT · WebVTT · TXT |
Last verified: September 2026. Results vary by hardware and audio quality.