Auto Subtitle Generator
Generate accurate subtitles from any video or audio file using AI speech recognition that runs entirely in your browser. Burn subtitles directly into your video with custom font, color, and placement — perfect for Reels, TikTok, and Shorts. Export as SRT or VTT.
Your media is processed in your browser and is not uploaded to our servers for conversion. Page assets (including the converter engine) still load over the network.
Select or drop a video or audio file
MP4, WebM, MOV, MP3, WAV, M4A, FLAC, OGG. The AI runs in your browser — nothing is uploaded.
Click, tap, or drop a file
Your media is processed in your browser and is not uploaded to our servers. The AI model downloads once to your browser cache and runs locally after that.
How to generate subtitles from a video
- Select a file. Drop or choose a video (MP4, WebM) or audio file (MP3, WAV, M4A, FLAC).
- Choose a language. Pick the spoken language for better accuracy, or leave it on "Auto-detect" to let the AI figure it out.
- Generate subtitles. Click Generate Subtitles. The AI model downloads on first use (~150 MB, cached for next time), then transcribes your audio.
- Edit the transcript. Review the timestamped segments and fix any words the AI got wrong. Click any segment text to edit it.
- Export or burn in. Download subtitles as SRT or VTT, or burn them directly into your video with custom font, color, and position — ready for social media.
How browser-based subtitle generation works
The tool uses OpenAI's Whisper speech recognition model compiled to run in your browser via ONNX Runtime WebAssembly. The model file (~150 MB) downloads on first use and is cached by your browser for instant startup on return visits.
Audio is extracted from your video using the Web Audio API, resampled to 16 kHz mono, and fed to the Whisper model in 30-second overlapping chunks. Each chunk produces timestamped text segments that become your subtitle lines.
Because everything runs locally, there are no usage limits, no per-minute charges, and no risk of your private audio being sent to a third-party transcription service.
Subtitle generator features
Runs entirely in your browser
The Whisper AI model downloads once and runs locally using WebAssembly. Your audio never leaves your device — not even for transcription.
Accurate AI speech recognition
Powered by OpenAI's Whisper model, the same architecture used by professional transcription services. Handles accents, background noise, and natural speech.
99 languages supported
Whisper supports transcription in English, Spanish, French, German, Japanese, Chinese, Hindi, Arabic, and 90+ other languages. Auto-detection works for most content.
Editable transcript
Every subtitle segment is editable before export. Fix names, technical terms, or any words the AI misheard — directly in the browser.
Burn subtitles into video
Embed subtitles permanently into your video with customizable font color, size, and placement (top, middle, or bottom). Download a single file ready for Instagram Reels, TikTok, YouTube Shorts, or any platform.
SRT and VTT export
Export as SubRip (.srt) for video editors and media players, or WebVTT (.vtt) for HTML5 video and web players. Both formats are industry standard.
Works with video and audio
Drop an MP4, WebM, MP3, WAV, M4A, FLAC, or OGG file. The tool extracts audio automatically from video files and transcribes it.
Frequently asked questions
Is my audio uploaded to a server for transcription?
No. The AI model runs in your browser using WebAssembly. Your audio stays on your device throughout the entire process. The only network request is to download the model file on first use.
How accurate are the subtitles?
Whisper is one of the most accurate open-source speech recognition models available. Accuracy depends on audio quality, background noise, and language. Clear speech in a quiet environment typically produces very accurate results. You can edit any mistakes before exporting.
What languages are supported?
Whisper supports 99 languages including English, Spanish, French, German, Italian, Portuguese, Japanese, Chinese, Korean, Hindi, Arabic, Russian, and many more. Use Auto-detect or select the language manually for better accuracy.
How long does transcription take?
Speed depends on your device and the audio length. On a modern laptop, expect roughly 1–3 minutes per minute of audio with the Tiny model. The first run is slower because the model needs to download (~150 MB). After that, it loads from cache almost instantly.
What is the difference between SRT and VTT?
SRT (SubRip) is the most widely supported subtitle format — it works in VLC, Premiere, DaVinci Resolve, and most video editors. VTT (WebVTT) is the web standard used in HTML5 video elements and streaming platforms. Both contain the same timed text; pick whichever your player or editor needs.
Can I edit the subtitles before downloading?
Yes. Every subtitle segment is editable. Click on any line to fix text, correct names, or adjust wording. Your edits are included in the exported SRT or VTT file.
Can I burn subtitles directly into the video?
Yes. After generating subtitles, choose your preferred font color, size, and position (top, middle, or bottom), then click 'Burn Subtitles into Video'. The tool re-encodes the video with hardcoded subtitles using FFmpeg WebAssembly — everything runs in your browser. The output is a single video file with subtitles permanently embedded, ready for social media.
Does it work offline?
After the AI model has downloaded once, the tool works offline. The model is cached in your browser, so subsequent visits do not require an internet connection.
What file formats can I use?
You can use video files (MP4, WebM, MOV) or audio files (MP3, WAV, M4A, FLAC, OGG, AAC). The tool automatically extracts the audio track from video files.
Related: Video trimmer · Video compressor · All tools