What Are Subtitles and How Do You Get Them for a Video?
Subtitles are timed text lines that appear on screen while a video plays, translating or transcribing the spoken audio so viewers can read along. You get them either by generating them automatically from the video's audio (with a tool like Scribe) or by writing and timing them manually. The right method depends on whether you need a quick readable transcript, a downloadable subtitle file, or a polished, perfectly synced track for publishing.
Subtitles vs. captions vs. transcripts
These three terms get used interchangeably, but they describe different outputs:
| Term | What it is | Typical use |
|---|---|---|
| Subtitles | Timed text shown on screen, often translating dialogue into another language | Watching foreign-language content, reaching multilingual audiences |
| Captions | Timed text that also conveys non-speech audio (sound effects, speaker labels) | Accessibility for deaf and hard-of-hearing viewers |
| Transcript | The full text of the spoken content, usually without timing | Reading, searching, repurposing into articles or notes |
In practice, the same source audio can produce all three. A tool that transcribes a video gives you the text; that text can then be formatted as subtitles or captions with timestamps attached.
Common subtitle formats and when to use them
Two formats cover most needs:
- SRT (SubRip Subtitle) — the most widely accepted format. Plain text with sequence numbers, start/end timestamps, and the subtitle line. Use it for YouTube uploads, most video editors, and players like VLC.
- VTT (WebVTT) — the web standard, used for HTML5 video and streaming platforms. It supports styling and positioning cues that SRT does not.
If you are unsure, SRT is the safer default because almost every platform imports it. Scribe's transcript panel offers Copy TXT SRT, so you can pull the plain text or the timed subtitle file directly from the same result.
How subtitles are created: automatic vs. manual
Automatic (auto-generated) subtitles come from speech-recognition software that listens to the audio and produces timed text. They are fast and cheap, and they work well when audio is clear and the speaker's language is well supported. Scribe supports 50+ languages and can pull from auto & manual captions, with instant language switching.
Manual subtitles are written or corrected by a person. They are slower and cost more, but they handle accents, overlapping speakers, jargon, and timing precision that automation can miss.
A practical middle path: generate automatically, then fix the errors that matter. One reviewer of Scribe noted they were "really looking forward to the punctuation feature," which is a fair reminder that auto output sometimes needs a cleanup pass before it is publish-ready.
How to get subtitles for a video
Using a transcription tool such as Scribe:
- Provide the video. Paste the video URL or use the browser extension on the page where the video plays. Scribe also offers a Chrome extension.
- Wait for transcription. The tool processes the audio and returns a timed transcript. The page shows a progress state (for example, "Transcribing... 29:45") while it works.
- Review the text. Read through for accuracy, punctuation, and speaker changes.
- Export in the format you need. Use the copy/download options to get TXT for a plain transcript or SRT for a subtitle file you can upload to a video platform.
The expected result is a text file you can either read directly or attach to your video as a subtitle track.
Multi-language subtitles and switching
If your audience spans languages, look for a tool that supports multiple subtitle languages and lets you switch between them. Scribe lists 50+ languages supported with instant language switching, which means one transcription can serve several language tracks rather than requiring a separate pass per language. This is useful for creators publishing the same video to different regional audiences.
Common issues to watch for
- Missing or wrong punctuation. Auto-generated text often runs together without commas or periods. Plan to edit before publishing.
- Sync problems. Subtitles that drift out of time with the audio usually mean the timestamps need adjusting, or the source audio had gaps or speed changes.
- Speaker confusion. Without speaker labels, multi-person dialogue can be hard to follow. Captions that identify speakers solve this.
- Language mismatches. If the detected language is wrong, the whole transcript will be off. Confirm the language setting before exporting.
Choosing an approach
- Need a quick readable transcript or a downloadable SRT? Automatic transcription is the fastest route, and a free tool covers most personal and creator use.
- Publishing to a broad or accessibility-focused audience? Budget time for a manual cleanup pass, or use captions that include non-speech audio.
- Working across languages? Prioritize a tool with multi-language support and easy switching so you are not re-transcribing per language.
Scribe reports being used by 550K+ creators with 1M+ videos transcribed, and offers a free Chrome extension alongside its web tool, so it is a reasonable starting point if you want to test automatic subtitles before committing to a paid workflow.