← all tools / audio & video

Subtitle Generator

Live

Create captions for video and get a timed SRT or VTT file for YouTube or a web player. No watermark, no signup.

What it does

  • SRT for YouTube and editors, VTT for an HTML5 video player
  • Timed from the speech itself, not a transcript with times pasted on
  • One video, up to 50 MB and up to 60 minutes long
  • Transcribed by our own faster-whisper worker, not a third-party AI service

Your files

Gone within the hour.

Files are uploaded over HTTPS, used only to run this tool, and deleted from our server automatically about an hour later. Never sold, never used to train anything.

Speech is processed by our self-hosted faster-whisper worker. Your recording is not sent to a third-party AI service.

Frequently asked questions

How do I generate subtitles from a video?

Drop the video in above, choose SRT or VTT, and press Generate. It is uploaded over HTTPS, transcribed on our server, and comes back as subtitles.srt or subtitles.vtt, which you can read on screen before you save it. A stretch of video with no speech in it produces no cue at all rather than an empty numbered one, so the cue numbers run unbroken from 1. An hour of video takes a while to caption and you do not have to sit and watch it: tick Play a sound when ready and a short chime tells you the file has arrived. There is no account to make, nothing to install, and no watermark on the file or the text.

What is the difference between SRT and VTT?

They carry the same cues and differ in the details a player cares about. SRT separates the milliseconds in a timecode with a comma (00:00:02,000) and VTT uses a full stop (00:00:02.000), which is the most common reason a subtitle file looks right and loads nowhere. VTT also begins with a WEBVTT line and treats its text as markup, so an ampersand or an angle bracket is escaped there and left exactly as spoken in SRT. One rule they share: a blank line inside a cue is removed before either file is written, because in SRT a blank line is what ends a cue, and one arriving mid-sentence would split that cue in two. Pick SRT for YouTube, Premiere and VLC, and VTT for an HTML5 track element on a web page.

How long and how large can the video be?

One video per job, up to 50 MB, and up to 60 minutes of running time. Common MP4, MOV, M4V, WebM, MKV, AVI, MPEG, OGV, FLV and transport stream files are accepted, and what decides that is the file rather than its name: the server reads the first four kilobytes and matches the container signature stored there, so a clip whose extension is wrong for what it actually contains is still recognised, and something that is not a video at all is turned away in the same breath instead of after a transcription attempt. The video also has to actually contain speech: one with none in it is refused with a message rather than handed back as an empty file.

Are the subtitles accurate enough to publish as they are?

Read them first. This is automatic speech recognition, so it hears names, jargon, accents and people talking over each other imperfectly, and we do not publish an accuracy percentage because the honest answer depends entirely on your audio. Two things help: choose the spoken language instead of leaving it on auto detect, and type the names and technical terms your video uses into the box before you start. The language list holds 18 beside Auto detect, and the server accepts exactly the same 18, so the dropdown can never offer you something the worker would then refuse. The terms box takes up to 400 characters and counts them under the box as you type, which makes it a place for the handful of proper nouns a recogniser cannot guess rather than a glossary. Treat the result as a good first draft, not a finished caption track.

Will the timings line up with my video?

Yes. The cues are timed from the same speech pass that produces the words, so nothing is estimated after the fact. They are sorted into order and clamped so that one cue can never overlap the next. A very short cue is stretched towards a full second so it is not gone before it can be read, though the clamp wins where there is no room for it: a 0.2 second line with silence after it becomes a full second, while one whose neighbour starts 0.5 seconds later stays at 0.5. Lines wrap at 42 characters, the width broadcast subtitle guidelines settled on, and a line that lands on exactly two is rebalanced into an even pair rather than left as one full line and a short one. No word is ever cut or dropped to make a line fit: a word longer than the limit keeps its own line.

Is it free, and what happens to my video?

It is free, with 5 speech jobs a day per connection, shared with Video to Text because both tools run the same pass. Busy does not mean turned away: if the speech worker is already occupied your video waits its turn, and the page shows how many are ahead of it and counts them down rather than leaving you at a spinner. Only once ten videos are waiting or running between the two tools does the site ask you to come back a little later instead of adding an eleventh. Your video is used only for this one job and deleted from our server automatically about an hour later, along with the subtitle file, so save your download before you close the tab. Speech is transcribed by our own faster-whisper worker running on our server and is never sent to a third-party AI service, sold, or used to train anything.

More Audio & Video tools

all tools