Audio to Text
Drop in a recording and a real Whisper AI model transcribes it, right here on your device. No upload, no sign-up, no minute limits. Get clean text plus timestamped SRT and VTT subtitles.
🔒 100% private — the model runs locally, your audio never leaves your deviceDrop an audio or video file here
or click to choose · MP3, WAV, M4A, OGG, WebM, MP4 · any length
How does this run locally?
This page runs OpenAI's Whisper speech-recognition model directly in your browser using transformers.js. When your device supports WebGPU it runs on the GPU; otherwise it falls back to WebAssembly on the CPU. The model downloads once (about 80 MB) and is then cached, so later files start instantly. Your audio is decoded and transcribed on your own machine, which is why there is no upload, no queue, and nothing private leaving your device.
Why on-device matters
Most transcription tools send your audio to a server, meter you by the minute, and hide the result behind a sign-up. Running the model locally removes all of that. Your recordings stay yours, it costs nothing per minute, and there is no account to create. It even keeps working offline once the model is cached, which is handy for interviews, lectures, and voice notes you would rather not hand to a third party.
Transcript and subtitles
You get an editable transcript plus a list of timestamped segments. Export plain text for notes, or SRT and VTT subtitle files to drop straight into a video editor or upload alongside a clip. Whisper is multilingual and detects the spoken language automatically, and you can switch the task to translate non-English audio into English text.
Is my audio uploaded anywhere?
No. The Whisper model is downloaded to your browser and every file is transcribed locally. Your audio never touches a server.
Is it free, and are there limits?
Completely free, no sign-up, no per-minute cap, no watermark. It runs on your machine, so there is no cost to pass on.
Why is the first run slow?
The first time, the model downloads (about 80 MB) and your browser caches it. After that it loads instantly and transcription is much faster, especially on a machine with WebGPU.
What can I export?
An editable text transcript, a plain .txt file, and SRT or VTT subtitle files with timestamps for video.
Does it work on phones?
It runs in modern mobile browsers, but transcription is heavy. A laptop or desktop, ideally with WebGPU, is faster and more reliable for longer recordings.
I build tools like this every day.
Senior full-stack engineer, available for senior or contract work, fully remote. See the rest of the lab or get in touch.