AnyFormat

00 VTT → TXT

VTT to TXT

Turns a caption track into prose you can actually read.

01VTT in
02TXT out
03Options 7

Paragraphs join consecutive cues into readable prose and start a new block at a pause.

Negative moves subtitles earlier. Use this when they run consistently ahead of or behind the audio.

For subtitles that drift progressively. 25 fps material played at 23.976 needs 1.0427.

Ends each cue just before the next begins, so no two are ever on screen at once.

  • No upload The conversion runs in this tab. Your file never travels — not to us, and not to the advertising.
  • No sign-up, no email, no daily cap There is no account system to sign up to.
  • No size limit we invented Only your device's memory — about ~2 GB on a desktop browser, ~400 MB on a phone.
  • Turn your Wi-Fi off and convert anyway The conversion needs nothing but this page. Ads will not load without a connection; your file will still convert. That is the whole claim, and it takes five seconds to check.

Turns a WebVTT caption track into readable text. Most auto-generated captions — from YouTube, Zoom, Teams or a transcription service — arrive as VTT, and they are close to unreadable as a cue list.

Auto-generated captions need more cleaning than authored ones

A machine transcript is cued every few words rather than by sentence, often with no punctuation at all, and frequently with overlapping cues that repeat words as the recogniser revises its guess. Joining cues into paragraphs helps a great deal; it cannot invent punctuation that was never there.

Cue settings, STYLE blocks and NOTE comments are all discarded — none of them is text.

Duplicated words from rolling captions

YouTube's live captions use a rolling window where each cue repeats the tail of the previous one. That repetition is in the source, not introduced here, and no converter can reliably remove it without guessing at meaning. If your transcript reads with every few words doubled, download the non-rolling caption track instead — YouTube offers both.

Speaker labels and positioning cues

WebVTT carries more than SRT does, and most of it is not dialogue. Voice spans mark who is talking as <v Alice>, cue settings put a caption in a particular corner, and NOTE blocks hold comments the viewer never sees. The positioning and the notes are dropped. Speaker names are worth keeping for a transcript, so they are preserved as a prefix rather than thrown away with the rest of the markup — a transcript that does not say who spoke is a much less useful document than one that does.

Other names for this

Also searched as “webvtt to text”, “youtube captions to text”, “vtt to transcript”.

Questions

Words are repeated throughout my transcript.
That is a rolling caption track, where each cue overlaps the last. Download the standard track rather than the live one.
There is no punctuation.
Then there was none in the captions. Automatic transcription often omits it entirely, and adding it would mean guessing at meaning.
Is my subtitle file uploaded?
No. The parsing and rewriting happen in this tab. Subtitle files often contain an unreleased script or a client’s content, which is a good reason not to send them to a stranger’s server. Turn your Wi-Fi off and the tool keeps working.
Can I convert several files at once?
Yes — drop as many as you like, or a whole folder. Each keeps its original name with the new extension, and you can download them individually or as a ZIP.