Video to Text
Turn video speech into a clean transcript
6 free tools for recordings: the words that were said, those words timed as subtitles, the sound on its own, or the same video in a file that is smaller or that actually plays. No account, no watermark, and no trial that expires halfway through a job.
Most people arrive here holding a recording and a deadline. Find the row that matches the problem in front of you.
| What you have | What you get | Tool |
|---|---|---|
| A meeting or lecture recording has to become notes | A transcript you can read, copy or download, with timestamps if you want them. | Video to Text |
| You want the music or the talk without the picture | The audio track on its own, either re-encoded to MP3 or lifted out untouched. | Video to Audio |
| A MOV from an iPhone or a Mac will not open on Windows | The same video inside an MP4, usually without re-encoding a single frame. | MOV to MP4 |
| A video is too large to email or upload somewhere | The same video re-encoded smaller, with the result measured against the original. | Compress Video |
| A video needs captions burned into a player or uploaded to YouTube | A timed SRT or VTT file, cut from the same speech pass that produces the words. | Subtitle Generator |
| You need a searchable record of what was said, not a file | The transcript is plain text, so it is searchable the moment you have it. | Video to Text |
This is the part most converters are vague about, so it is worth saying plainly. Re-encoding a video or a piece of audio always costs quality, because it decodes what is there and compresses it again. Two of the tools here avoid that step entirely whenever they can, and Compress Video is the one that cannot: reducing a file size is re-encoding, which is why its result screen reports what the trade actually cost rather than leaving you to guess.
MOV and MP4 are containers: different boxes around the same encoded video. A MOV from an iPhone, a Mac screen recording or most cameras is already carrying video that MP4 holds natively, so MOV to MP4 moves the streams across without decoding them and the picture that comes back is identical. Only footage ordinary players cannot open, ProRes from an editing program being the usual case, has to be encoded again. The result screen tells you which of the two happened rather than letting you assume the good one.
Video to Audio offers the same deal on sound. Its default is an MP3, which is a second lossy pass over audio that a video had already compressed once, and no bitrate recovers detail the recording never held. The alternative copies the existing track into its own file without decoding a sample, so the sound is exactly what the video carried. It is also the fastest option and usually the smallest file. When a track has no container it can sit in on its own, the job is refused with a message instead of quietly re-encoding it and calling that the same thing.
Every tool here takes one file of up to 50 MB. After that they part company, and a page that quoted one set of numbers for the category would be wrong about most of it.
| Tool | Longest video | Largest result | Daily allowance |
|---|---|---|---|
| Video to Text | 60 minutes | Plain text | 5 a day shared, 10 at once |
| Subtitle Generator | 60 minutes | SRT or VTT | 5 a day shared, 10 at once |
| Video to Audio | 60 minutes | 250 MB | No |
| MOV to MP4 | 20 minutes | 500 MB | No |
| Compress Video | 20 minutes | 500 MB | No |
Speech is the only rationed part, and the reason is worth knowing: it runs a model that holds a worker for as long as the audio lasts, while the conversion tools are ffmpeg finishing in far less than real time. So it gets five videos a day per connection and a queue that admits ten at a time. That allowance is one shared pool, not five each: a transcript and a subtitle file are the same speech pass with a different renderer at the end, and they post to the same endpoint, so five subtitle jobs use up all five for the day. Reach either limit and you get a message saying so. Going over a length or result limit is a refusal with the number in it too, never a job that runs for ten minutes and then fails.
Turn video speech into a clean transcript
SRT and VTT subtitles from any video
SRT to VTT and back, in your browser
Extract the MP3 from a video
Make a QuickTime MOV play everywhere
Make a big video smaller
In the workshop
Fast, accurate speech-to-text
Clip a moment into a GIF
Natural voices, free
Extract an MP3 audio file from an MP4 video
Cut a video to the exact section you need
Convert audio between common formats
Cut an audio file without complicated editing
Turn animated GIF files into MP4 videos
Convert MKV videos into MP4 files
Convert WebM videos into MP4 files
Remove the audio track from a video
Speed up or slow down a video
Make quiet audio louder
Join multiple audio files in order
Record and save audio in your browser
Record a screen, window, or browser tab
Move every subtitle forward or backward in time
Your file is uploaded
Unlike the quick utilities on this site, these cannot run in your browser: a video is too large for that. The file goes to our server over HTTPS, is used for your one job, and is deleted with whatever it produced about an hour later. Subtitle Converter is the one exception: it only rewrites a small text file, so it runs in your browser and uploads nothing.
One file at a time
Each of these takes a single video per job, up to 50 MB. There is no batch mode and no folder upload, which is the honest limit of a browser tool rather than something a paid tier unlocks.
Speech stays here
Transcription runs on a faster-whisper worker on our own machine, so a recording is not handed to an outside AI service. None of the others involve a model at all.
Tags are left behind
An MP4 or an extracted audio file comes back without the container tags the original carried: the device name, the software that recorded it, and any location stored beside the video stay with the file you uploaded.
A 50 MB ceiling rules out most footage straight off a camera, and no option here raises it. A feature-length recording, a folder of clips, or anything you will do again next week wants a local copy of ffmpeg rather than a browser tab: it is free, it has no size limit, and it will be faster than uploading.
Two more cases belong elsewhere. A transcript that has to be right rather than roughly right still needs a human pass, because automatic speech recognition guesses at names, crosstalk and anything competing with background noise, and it guesses confidently. And a recording that is confidential in the way a legal or medical one is confidential deserves the tool your workplace has approved: our deleting it an hour later is easy for us to state and hard for you to verify from outside, which is exactly why the mechanism is written down above rather than asserted.