SRT vs VTT: Subtitle File Formats Explained (and How to Convert Them)

SRT (SubRip) is a plain-text subtitle format with numbered cues and comma-separated timestamps. WebVTT (.vtt) is the web's subtitle format, with a WEBVTT header, dot-separated timestamps and support for positioning, styling and speaker labels. Use SRT for the widest compatibility and VTT for HTML5 players. Converting is easy, but VTT-only features can be lost.
Subtitle files look alike and break in different ways. This guide explains how SRT and WebVTT are built, how they differ, which platforms accept which, how to convert between them (including what gets lost) and the readability limits broadcasters use. For what captions, subtitles and SDH mean, start with Captions vs Subtitles vs SDH; for generating the captions in the first place, see how to add auto captions to any video.
Facts checked October 9, 2026 against the W3C WebVTT specification, the HTML Standard, the YouTube, Facebook, LinkedIn, Vimeo and Brightcove help pages, the Library of Congress format record, FFmpeg's source code and the Netflix and BBC subtitle guidelines. Full disclosure: Loopdesk is our product and makes the free SRT to VTT converter linked here.
What is an SRT file?
An SRT file is a plain-text subtitle file made of numbered blocks: a counter, a start --> end timecode with a comma before the milliseconds, one or more lines of text, and a blank line. It has no formal standard; the Library of Congress calls it "a very simple text format, but not well standardized."
1
00:00:01,000 --> 00:00:03,400
Welcome back to the show.
2
00:00:03,600 --> 00:00:07,200
Today we're talking about AI video editing
and why captions matter.
- Tags: many players understand bold, italic and underline tags and some support font color, but support varies by player.
- Encoding: the format defines none (SubRip's default was Windows-1252), so platforms set their own rule, and modern ones expect UTF-8. YouTube says SRT "must be in plain UTF-8."
What is a VTT (WebVTT) file?
WebVTT is a text-track format from the W3C, designed for the HTML track element. A file starts with the WEBVTT signature, uses a period before the milliseconds, may omit the hour, and can carry cue settings (position, alignment, size), STYLE blocks styled with CSS, NOTE comments and speaker voice tags.
WEBVTT
NOTE Cue settings such as align and position are optional
intro
00:00:01.000 --> 00:00:03.400 align:start position:10%
<v Host>Welcome back to the show.
00:00:03.600 --> 00:00:07.200
<v Host>Today we're talking about AI video editing
and why captions matter.
The W3C specification is a Candidate Recommendation Draft dated May 20, 2026. It has not been published as a final Recommendation, and W3C marks it a work in progress. Its MIME type is text/vtt, registered with IANA in 2019. In the example, intro is an optional cue identifier, align:start position:10% are cue settings, and <v Host> is a voice tag naming the speaker.
SRT vs VTT: what is the difference?
The main differences are the header, the timestamp separator and what each can express. SRT has no header, uses commas (00:00:01,000) and supports only basic tags. WebVTT starts with WEBVTT, uses periods (00:00:01.000) and adds positioning, CSS styling, voice tags, comments and chapter or metadata tracks. SRT is simpler; VTT is richer and web-native.
| Feature | SRT | WebVTT |
|---|---|---|
| Standard status | No formal standard (Library of Congress: "not well standardized") | W3C Candidate Recommendation Draft (May 20, 2026) |
| First line | None (it starts with cue number 1) | WEBVTT |
| Timestamp | 00:00:01,000 (hours required) | 00:00:01.000 (hours optional: 00:01.000) |
| Cue identifiers | Sequential numbers, by convention | Optional; unique if used |
| Styling | Basic tags (bold, italic, underline; sometimes font color), player-dependent | Cue tags plus CSS through STYLE blocks and ::cue |
| Positioning | None | Cue settings (position, align, size, line, vertical) and regions |
| Speaker labels | None (write "HOST:" in the text) | Voice tags: <v Name> |
| Comments | None | NOTE blocks |
| Chapters and metadata | No | Yes, through <track kind="chapters"> and kind="metadata" |
| Encoding and type | Not defined; UTF-8 expected by platforms | UTF-8, text/vtt |
HTML5 <track> element | Not natively; the track loader parses WebVTT | Yes |
Which format does each platform accept?
It depends on the platform. YouTube accepts both, plus a dozen or so more; Facebook and LinkedIn accept SRT only; Vimeo accepts SRT and WebVTT and recommends WebVTT; HTML5 players need WebVTT. Instagram and TikTok caption-file uploads aren't documented on the official pages we could read, so burn captions in there.
| Platform | Accepted caption files | Details (official pages, checked October 9, 2026) |
|---|---|---|
| YouTube | .srt, .sbv/.sub, .mpsub, .lrc, .cap, .smi/.sami, .rt, .vtt, .ttml, .dfxp, .scc, .stl, .tds, .cin, .asc | SRT: "No style info (markup) is recognized. The file must be in plain UTF-8." VTT: "Positioning is supported, but styling is limited to <b>, <i>, <u>." You can upload with timing or without timing (YouTube syncs a plain transcript); transcripts aren't recommended for videos over an hour or with poor audio |
| SRT only | Name the file filename.[language code]_[country code].srt, for example filename.en_US.srt (lowercase language, uppercase country); hh:mm:ss,fff timestamps; unique cue numbers starting at 1; UTF-8. The Graph API limits caption files to 200K | |
| SRT | "You must have an associated SRT (SubRip Subtitle) file attached to the video before it can be posted." Desktop upload. The help page shows "last updated 5 years ago", so recheck in the app | |
| Vimeo | SRT and WebVTT | "We recommend using WebVTT." UTF-8 required |
HTML5 <track> | WebVTT | The HTML Standard defines text-track loading for WebVTT; other formats fail to load |
| Brightcove | Several | Converts non-WebVTT formats to WebVTT; UTF-8 mandatory |
| Instagram, TikTok | Not documented on the official pages we could read | Burn captions into the video |
How do you convert SRT to VTT (and back)?
To convert SRT to VTT, add a WEBVTT line and a blank line at the top, change the comma before each millisecond value to a period, and save as UTF-8 with a .vtt extension. To go back, remove the header and notes, restore commas, number the cues, and accept that styling and positioning will be lost.
Three ways to do it:
-
Loopdesk's free converter. The SRT to VTT converter runs in your browser. Per the page, it converts SRT to VTT, VTT to SRT and subtitles to plain text, with no upload, no signup and no watermark, and it can shift timing (cues that would end before 0:00 are dropped).
-
FFmpeg. One command each way. FFmpeg writes VTT timestamps without the hour when it is zero (
00:01.000), which is valid WebVTT, and drops the cue numbers.ffmpeg -i captions.srt captions.vtt ffmpeg -i captions.vtt captions.srt -
Find and replace, for a tiny file. This adds the header and swaps the separator:
{ printf 'WEBVTT\n\n'; sed -E 's/([0-9]{2}:[0-9]{2}:[0-9]{2}),([0-9]{3})/\1.\2/g' captions.srt; } > captions.vtt
Conversion checklist:
- Encoding. Save as UTF-8. WebVTT allows an optional byte-order mark before
WEBVTT. - Escaping. WebVTT cue text may not contain a raw
&or<except as part of a character reference or tag, so write&and<. - Identifiers. Cue identifiers must be unique in VTT, and cue text can't contain
-->. - Overlaps. WebVTT allows overlapping cues, but players differ in how they stack them.
- Test in the target player. A file that validates can still look wrong on one platform.
What gets lost when you convert VTT to SRT?
Anything SRT can't express: cue settings such as position and alignment, regions, STYLE blocks and CSS, NOTE comments, speaker voice tags, class and language tags, inline timestamps and chapter cues. Many converters keep only bold, italic and underline. FFmpeg's WebVTT decoder, for example, maps just those three tags and skips STYLE, REGION and NOTE blocks.
Loopdesk's free converter keeps bold, italic and underline, writes <v Name> voice tags as Name: in SRT and plain-text output, escapes a raw & or < when converting SRT to VTT, and drops positioning, STYLE and NOTE blocks.
Keep the VTT as your master if you use positioning or speaker labels, and export SRT only for platforms that require it. Speaker identification matters for accessibility, so if a converter drops <v Name> tags, add the names back as text (for example HOST: …) before you deliver the file. Our guide to Are Auto Captions ADA and WCAG Compliant? explains why.
What about SBV, ASS, TTML, SCC and STL?
Beyond SRT and VTT you'll meet SBV (a basic SubViewer-style format YouTube accepts), ASS/SSA (advanced styling and karaoke), TTML and IMSC (W3C XML formats used in broadcast and streaming delivery), SCC (CEA-608 closed captions) and EBU-STL (a European broadcast exchange format). Most creators only need SRT or VTT; the rest matter for broadcast or player-specific work.
| Format | What it is | Source note |
|---|---|---|
| SBV (.sbv, .sub) | SubViewer-style text | YouTube: "Only basic versions of these files are supported." |
| ASS / SSA | Advanced SubStation Alpha | Matroska: "advanced display features, like positioning, karaoke, or style managements" |
| TTML / IMSC | W3C XML formats for timed text | TTML2 is a W3C Recommendation (November 8, 2018); IMSC 1.2 is a Recommendation (August 4, 2020) |
| SCC | Scenarist Closed Caption | YouTube: "exact representation of CEA-608 data" |
| EBU-STL | European broadcast subtitle exchange format | EBU Tech 3264-E (February 1991) |
What line length and reading speed should subtitles follow?
Netflix's English (USA) guide sets 42 characters per line, a maximum of two lines and up to 20 characters per second for adult programs; the BBC recommends 160-180 words per minute. These are broadcaster guidelines, not platform rules, but they are a sound baseline for readable captions on any video.
Netflix's general requirements also cap a single subtitle event at 7 seconds, and its limits apply to English (USA); other languages have their own guides. A quick quality check before you deliver a file:
- UTF-8, with accents and non-Latin characters displaying correctly.
- No empty cues, and no overlapping cues unless you meant them.
- Two lines at most, broken at natural phrases.
- Names and brand terms spelled consistently.
- The last cue ends before the video does.
- Played back once in the target platform or player.
Should you upload a caption file or burn captions in?
Upload a caption file (SRT or VTT) wherever the platform supports caption tracks, so viewers can turn captions on or off and screen readers can read them. Burn captions in for Shorts, Reels and TikTok, where style travels with the video. Many creators publish both: a styled burned-in version and a sidecar file for platforms that accept one.
In Loopdesk: Aura generates automatic captions in 108 languages, and you can review and edit them before export. Export a video with burned-in captions, or SRT and VTT sidecar files for platform-native captioning. Captions and transcripts also feed discoverability; see Video SEO in 2026.
Sources
Read and verified on October 9, 2026.
- W3C, "WebVTT: The Web Video Text Tracks Format" (Candidate Recommendation Draft, May 20, 2026): https://www.w3.org/TR/webvtt1/ (dated version: https://www.w3.org/TR/2026/CRD-webvtt1-20260520/)
- IANA, media type registration for
text/vtt: https://www.iana.org/assignments/media-types/text/vtt - WHATWG, HTML Standard, media elements and the track element: https://html.spec.whatwg.org/multipage/media.html
- Library of Congress, "SubRip Subtitle format (SRT)": https://www.loc.gov/preservation/digital/formats/fdd/fdd000569.shtml
- Matroska, "Subtitles": https://www.matroska.org/technical/subtitles.html
- YouTube Help, "Supported subtitle and closed caption files": https://support.google.com/youtube/answer/2734698 · "Add subtitles & captions": https://support.google.com/youtube/answer/2734796
- Facebook Help (caption files): https://www.facebook.com/help/261764017354370 · Graph API, video captions: https://developers.facebook.com/docs/graph-api/reference/video/captions/
- LinkedIn Help, "Add Closed Captions to Videos on LinkedIn": https://www.linkedin.com/help/linkedin/answer/a552177/add-closed-captions-to-videos-on-linkedin
- Vimeo Help, "How to add captions or subtitles to my video": https://help.vimeo.com/hc/en-us/articles/21956884955537-How-to-add-captions-or-subtitles-to-my-video
- Brightcove, "Adding Captions to Videos using the Media Module": https://studio.support.brightcove.com/media/captions/adding-captions-videos-using-media-module.html
- FFmpeg source: https://github.com/FFmpeg/FFmpeg (libavcodec/webvttdec.c, libavformat/webvttdec.c, libavcodec/srtenc.c)
- W3C TTML2: https://www.w3.org/TR/ttml2/ · IMSC 1.2: https://www.w3.org/TR/ttml-imsc1.2/
- EBU Tech 3264-E: https://tech.ebu.ch/docs/tech/tech3264.pdf
- Netflix Partner Help, "English (USA) Timed Text Style Guide": https://partnerhelp.netflixstudios.com/hc/en-us/articles/217350977-English-USA-Timed-Text-Style-Guide · "Timed Text Style Guide: General Requirements": https://partnerhelp.netflixstudios.com/hc/en-us/articles/215758617-Timed-Text-Style-Guide-General-Requirements
- BBC, "BBC Subtitle Guidelines" (v1.2.5, March 2026): https://www.bbc.co.uk/accessibility/forproducts/guides/subtitles/
- Loopdesk, "Free SRT to VTT Converter": https://loopdesk.ai/tools/srt-converter
We could not load Zoom's or Wistia's help pages, or Instagram's and TikTok's caption-upload documentation, on October 9, so we make no claim about them.
Frequently Asked Questions
Is VTT better than SRT?
Neither is universally better. SRT is simpler and accepted almost everywhere captions can be uploaded. WebVTT is the W3C's text-track format for HTML5 players and supports positioning, styling, speaker labels and chapters. Vimeo recommends WebVTT; Facebook and LinkedIn accept SRT only.
Does YouTube accept VTT files?
Yes. YouTube's supported-files list includes .vtt along with .srt and about a dozen other formats. For VTT, positioning is supported but styling is limited to bold, italic and underline; for SRT, no style markup is recognized, and files must be plain UTF-8.
Can I use an SRT file on my website's video player?
Not directly with the HTML track element, which loads WebVTT. Convert the SRT to VTT by adding the WEBVTT header and changing commas to periods, then reference it with a track element. Some player libraries accept SRT and convert it for you, so check your player's documentation.
How do I convert SRT to VTT for free?
Use Loopdesk's free SRT to VTT converter in your browser: paste or upload the file, choose WebVTT, then copy or download the result. Per the tool page, nothing is uploaded and there is no signup or watermark. FFmpeg also converts between the two formats with a single command.
Do SRT files support styling?
Only minimally. Many players understand bold, italic and underline tags, and some support font color, but support varies by player, and YouTube recognizes no style markup in SRT files. For positioning and CSS styling, use WebVTT, or burn the captions into the video.
What character encoding should subtitle files use?
UTF-8. YouTube requires plain UTF-8 for SRT, Facebook and Vimeo specify UTF-8, Brightcove makes it mandatory, and WebVTT is defined as UTF-8. Older SRT files saved as Windows-1252 can show broken accents or non-Latin characters until you re-save them as UTF-8.
What is the difference between SRT and SDH?
SRT is a file format; SDH (subtitles for the deaf and hard of hearing) is a type of subtitle content that adds speaker identification and sound descriptions. An SDH track can be delivered as an SRT or VTT file. See our guide to captions, subtitles and SDH.
Have an SRT or VTT that needs converting, or a video that needs captions in the first place? Use the free converter, or open Loopdesk and ask Aura to caption the video in 108 languages. Plans start at $19/month ($15 billed annually); exports are never watermarked.