MT Media Tools

Transcribe Audio to Text Online — Private Whisper AI

Convert speech in audio and video files to text with Whisper AI running in your browser. No uploads, free, with copy and TXT download.

🔒 Runs entirely in your browser — nothing is uploaded

Drop an audio or video file here or click to browse

No file selected.

Advertisement

Speech to text without uploading your recordings

This transcriber turns spoken words in an audio or video file into plain text using OpenAI's Whisper speech recognition model. Unlike most online transcription services, it does not send your file anywhere. The model runs inside your browser through Transformers.js and ONNX Runtime, using your graphics card via WebGPU when it is available and falling back to WebAssembly on the CPU otherwise. Interviews, meeting recordings, lectures, voice memos and private videos stay on your device the whole time.

The only network download is the model itself. On the first run the tool fetches the Whisper weights from the Hugging Face model hub, which the browser then caches, so the next transcription starts almost immediately. Nothing about your audio is included in that request.

How the transcription works

When you press Transcribe, the audio track is decoded and resampled to 16 kHz mono, the format Whisper expects. Long recordings are split into windows of up to 30 seconds, and each cut is placed at the quietest moment near the end of the window so words are rarely chopped in half. Silent windows are skipped because Whisper tends to invent phrases such as "Thank you" when it hears nothing. Each window is transcribed in a background worker so the page stays responsive, and the text appears as soon as each piece is finished.

Whisper Tiny is the fastest option and works well for clear speech. Whisper Base is noticeably more accurate with accents and noisy audio at the cost of a larger download and slower processing. If you know the spoken language, select it; auto-detect listens to the first part of the recording and picks the most likely language. The Translate option produces English text from speech in another language.

Tips for better results

Speed depends on your hardware: with WebGPU a few minutes of audio often takes well under a minute, while CPU-only processing can take close to real time. Recordings with one speaker close to the microphone give the best text. Background music, crosstalk and heavy compression reduce accuracy. The output has no speaker labels or timestamps, so if you need captions with timing, use the auto subtitle generator instead. Always proofread names, numbers and specialist vocabulary before publishing or quoting the transcript. Very long files need a lot of memory; if a tab runs out of memory, split the recording into shorter parts first.

How to use

  1. Add a fileDrop an audio recording or a video with speech into the box.
  2. Pick model and languageChoose Tiny for speed or Base for accuracy, and set the spoken language or leave Auto-detect.
  3. TranscribePress Transcribe and watch the text appear window by window as the model works.
  4. Copy or downloadProofread the transcript, then copy it or save it as a .txt file.

Frequently asked questions

Is my audio uploaded to a server?
No. Your audio or video is decoded and transcribed on your own device. The only thing downloaded is the Whisper model itself, which comes from the Hugging Face model hub and is cached by your browser.
Why is the first transcription slower?
The first run downloads the speech recognition model (roughly 40–130 MB depending on the model you pick) and the ONNX runtime. After that they are cached, so later runs start much faster.
Which languages are supported?
Whisper understands about 100 languages. Choose the spoken language for the best accuracy, or use Auto-detect. You can also translate speech from other languages into English text.
How accurate is it?
Clear speech with little background noise transcribes well, especially with the Base model. Accents, music, overlapping speakers and technical terms can cause mistakes, so always proofread the text.
Does it work with video files?
Yes. MP4, WebM, MOV, MKV and most audio formats work. The audio track is extracted in the browser, with ffmpeg.wasm as a fallback for formats the browser cannot decode itself.
Advertisement