← all tools / audio & video

Video to Text

Live

Transcribe video to text online: upload a lecture, meeting, interview, or clip and get a transcript to read, copy, and download, in 18 languages.

What it does

  • Handles common phone, camera, and screen recording formats
  • Auto language detection with optional timestamped sections
  • Copy in the browser or download a plain text file

Your files

Gone within the hour.

Files are uploaded over HTTPS, used only to run this tool, and deleted from our server automatically about an hour later. Never sold, never used to train anything.

Speech is processed by our self-hosted faster-whisper worker. Your recording is not sent to a third-party AI service.

Frequently asked questions

How do I transcribe video to text online?

Upload one video, choose its language or leave Auto detect selected, then press Transcribe video. It runs online in your browser with no account and nothing to install. While the job runs the page reports the stage it has reached and an estimated time remaining, refined from the speed measured on your own file rather than a fixed guess, so a long recording tells you where it is instead of showing a spinner. If you would rather not sit and watch it, tick Play a sound when ready and a short chime tells you the transcript has arrived. You can read and copy the result in the page or download it as transcript.txt.

Which video formats can I transcribe?

The converter accepts common MP4, MOV, M4V, WebM, MKV, AVI, MPEG, OGV, FLV, and transport stream files. The video must contain a readable audio track: one with no audio stream at all is refused when its length is checked, and a video that has audio but no speech in it comes back as a message saying so rather than as an empty transcript.

What are the video size and length limits?

One video per job, up to 50 MB, and up to 60 minutes of running time. The length is read in the browser from the video itself, so a file that is too long is refused before the upload starts rather than after you have waited for it, and the server checks the real length again when it opens the file. The current free allowance is five speech jobs a day per connection, and it is shared with Subtitle Generator because both tools run the same transcription pass on the same endpoint: five between the two, not five each. If the speech worker is already busy your video waits in line rather than failing, and the page shows its place in the queue and counts it down as the videos ahead of it finish. Only when ten videos are already waiting or running does the site ask you to come back a little later instead of adding an eleventh.

Can the transcript include timestamps?

Yes. Keep Include timestamps selected and every spoken section starts with a marker in [HH:MM:SS] form, counted from the start of the video, one section to a line. Turn it off and the same sections are joined into continuous plain text. The markers are written into the text file you download, so they survive being pasted somewhere else. If you want a subtitle file with the timings kept as cues, use Subtitle Generator instead: it is the same speech pass written out as SRT or VTT.

Which languages can it transcribe, and how do I get a better result?

Eighteen, and Auto detect if you would rather not choose: English, Spanish, French, German, Italian, Portuguese, Dutch, Polish, Russian, Ukrainian, Turkish, Arabic, Hindi, Urdu, Chinese, Japanese, Korean and Vietnamese. Naming the language beats leaving it on Auto detect when you already know it, because detection is one more thing that can go wrong. Then use the Names or special terms box, which holds up to 400 characters, for the people, brands and technical words the recording uses, so the model has them before it starts rather than guessing at them. Clear speech with little background noise produces the best result.

What happens to my video after transcription?

Your video, extracted audio, and transcript stay on TryDeputize infrastructure and are deleted automatically within about an hour. A video waiting in the transcription queue is kept only until it runs, with a 12-hour safety limit for abandoned work. Speech is processed by our self-hosted faster-whisper worker, so your recording is not sent to a third-party AI service. That worker has a queue of its own, kept apart from the one the PDF and image tools use, so a long transcription never slows those down and a rush on them never delays yours.

More Audio & Video tools

all tools