Runs on your device — nothing is sent to a server
Transcribe audio to text
Drop an audio file and get a timestamped transcript. Whisper runs inside this tab — the recording never leaves your computer, and there is no account, queue, or monthly minute count.
Transcription controls
Drop an audio file anywhere on this page
Or paste one with ⌘V, pick one from disk, or record straight from your microphone.
mp3 · m4a · wav · ogg · flac · webm · mp4 (audio track) — up to 100 MB and 30 minutes per file
No file handy? Try:
checking this browser…
Transcript
Timestamps are hidden in this view. The .srt and .vtt downloads still carry them, and turning the switch back on brings the clickable segments back without re-running anything.
How it works
Pick the recording
Drag it in, paste it, choose it from disk, or hit Record. The file is read by this page only — there is no upload step to wait for.
The model downloads once
About 278 MB of Whisper weights land in the browser cache the first time. Every visit after that starts in under a second, even offline.
Your machine does the work
The audio is cut into 30-second blocks and transcribed block by block on your GPU or CPU, with the text appearing as each block finishes.
What you get
| Per file | Up to 30 minutes and 100 MB, one file at a time. |
|---|---|
| Accepted files | Anything your browser can decode: mp3, m4a, wav, ogg, flac, webm, and mp4 (the audio track is used). |
| Model | Whisper base — 74 million parameters, 99 languages, the same weights OpenAI released. |
| First-run download | 278 MB on WebGPU, 197 MB on the CPU backend. Cached afterwards. |
| Speed | On an Apple M2 a 30-second clip takes about 1.6 s on WebGPU and about 13 s on the single-threaded CPU backend; 10 minutes of audio lands near 40 s and just under 4 minutes respectively. |
| Output | Clickable timestamped segments, plus .txt, .srt and .vtt downloads and a one-click copy. |
Everything above is measured on this page with the sample clip, not copied from a model card. Longer recordings, more languages at once and speaker labels are what the Mac app is for.
Model and license
This page runs openai/whisper-base, converted to ONNX by onnx-community and executed with Transformers.js on ONNX Runtime Web. The weights are published under Apache-2.0, which permits commercial use; the original Whisper source code from OpenAI is MIT. Both licenses are reproduced in full, alongside the licenses for Transformers.js and ONNX Runtime, in the file below.
Model repository on Hugging Face Full license texts (LICENSE-model.txt)
Questions people actually ask
Does my audio get uploaded anywhere?
No. The file is read by JavaScript in this tab and decoded with the browser's own audio decoder; the samples are handed to a Web Worker running on your machine. The only things this page fetches over the network are the page itself and the model weights. You can prove it: open DevTools, switch to the Network tab, and run a transcription — after the weights are cached you can turn off Wi-Fi entirely and it still works.
Which languages does it handle?
Whisper base knows 99 languages. If you leave the selector on Auto-detect, the page runs one extra decoding step on the first 30 seconds to read the language token the model predicts, then transcribes in that language. English is what it does best. Chinese, Japanese, Korean and the smaller European languages work but make noticeably more mistakes at this model size.
How accurate is this compared with Otter or a paid service?
Honestly: worse. Paid services run models ten to twenty times this size and often add a language model pass on top. Whisper base gets clear single-speaker English close to right and will still mangle proper nouns — in the sample clip on this page it hears the podcast name "Half Duplex" as "Half-Dukeplex". It also cannot tell you who is speaking. What you get instead is a transcript of a recording you were never allowed to upload in the first place, for free, with no length quota.
Why does it download 278 MB the first time?
Because the model runs here rather than on a server, your browser needs the actual weights: a 78 MB encoder and a 200 MB decoder. They go into the browser's Cache Storage, so the second visit loads in well under a second and works with the network off. The CPU backend uses a 4-bit quantized decoder instead and only needs 197 MB. Clearing site data is what removes them.
My browser says CPU mode. Will it still work?
Yes, on WebAssembly running on a single CPU thread — that is the fallback for Firefox and Safari and for machines whose GPU driver WebGPU refuses. Expect a little under four minutes for a ten-minute recording instead of about forty seconds. If that matters, Chrome or Edge on the same machine will usually give you the WebGPU path.
Can I transcribe a video file?
If the browser can decode it, yes — drop an .mp4 or .webm and only the audio track is read. Very large video containers are the usual problem: the 100 MB limit counts the whole file, video included, so exporting audio-only first gets you far more minutes out of the same budget.
Longer recordings, and who said what
Whisper cannot separate speakers, and half-hour meetings are near the edge of what a browser tab should be asked to do. The Mac app handles both: speaker diarization, multi-hour recordings, and batch folders.
Open the Mac app