Turn a recording into text with AI speech recognition that runs entirely in your browser. Your audio is never uploaded, which makes it suitable for meetings, interviews and personal voice notes.
Your media is processed in your browser and is not uploaded to our servers for conversion. Page assets (including the converter engine) still load over the network.
Select or drop an audio or video file
MP3, M4A, WAV, OGG, MP4, MOV, WebM and more. Transcribed on your device.
Click, tap, or drop a fileChoose the recording
Drop or select an audio file (MP3, M4A, WAV, OGG) or a video (MP4, MOV, WebM).
Set the language
Pick the spoken language for better accuracy, or leave it on automatic detection.
Transcribe
The first time, the speech model (about 60 MB) downloads and is cached. Then transcription runs on your device.
Copy or download
Copy the text, or download it as plain text, timestamped text, SRT or VTT.
The tool runs Whisper, OpenAI's open-source speech recognition model, inside your browser using WebAssembly. The model file is downloaded from Hugging Face the first time and stored in your browser's cache; your recording is decoded and transcribed locally and is never sent anywhere.
It uses the compact 'tiny' version of Whisper so it can run on ordinary laptops and phones. It handles clear speech well and adds punctuation, but it makes more mistakes than large cloud services with background noise, strong accents, overlapping speakers or specialist vocabulary. It doesn't label who is speaking. Always review the transcript before relying on it.
The recording is processed on your device, so confidential conversations stay confidential.
No account, minutes quota or subscription.
Readable paragraphs, text with timestamps, or SRT and VTT subtitle files.
Once the model has downloaded, transcription doesn't need an internet connection.
No. Only the speech model is downloaded, once, from Hugging Face. Your recording is transcribed in your browser and never leaves your device.
Good for clear speech in a quiet setting. Accuracy drops with background noise, music, heavy accents, people talking over each other and technical terms or names. Choosing the spoken language helps.
It depends on your device. Short voice notes take seconds; an hour-long meeting can take several minutes on a laptop and longer on a phone. Keep the tab open while it works.
Up to 2 hours and 500 MB on a computer, or 1 hour and 200 MB on a phone, because the audio is decoded in memory. For a large video, convert it to M4A with the Audio Converter first; the audio alone is much smaller.
Over 20 are listed, including English, Hindi, Spanish, French, German, Portuguese, Arabic, Japanese, Chinese, Bengali, Tamil, Telugu, Marathi and Urdu. English gives the most accurate results with this model.
No. The transcript is one continuous text. Timestamped text makes it easier to find who said what.
Related: Auto Subtitles · Audio Converter · Audio Cutter · Increase Video Volume · All tools