How you-get Extracts and Saves Subtitles in SRT Format
you-get extracts subtitles by parsing XML caption tracks from video sources like YouTube, converting them into properly formatted SRT strings with calculated timestamps, and writing them to .srt files alongside the downloaded media through the base Extractor class.
The open-source video downloader you-get automates subtitle extraction across multiple platforms by standardizing disparate caption formats into SubRip (SRT). When you enable caption downloading via the --caption flag, specialized extractors parse platform-specific XML data and populate self.caption_tracks with SRT-formatted strings, while the generic base class in src/you_get/extractor.py handles file persistence. This workflow ensures consistent subtitle output regardless of whether the source is YouTube, iQiyi, or Acfun.
Parsing XML Captions into SRT Format
Platform-specific extractors handle the initial subtitle extraction. In src/you_get/extractors/youtube.py (lines 94-112), the YouTube extractor retrieves caption tracks from the ytInitialPlayerResponse['captions']['playerCaptionsTracklistRenderer']['captionTracks'] JSON path, downloads the corresponding XML files, and converts them to SRT format.
The conversion process iterates over <text> nodes from parseString(get_content(ttsurl)) and constructs valid SRT entries:
- Timestamp calculation: Each node provides a
startattribute and optionaldur(duration) attribute. The code calculatesfinish = start + dur(defaulting duration to 1.0 seconds if absent). - SRT time formatting: Seconds are formatted as
{:0>2}:{:0>2}:{:06.3f}and the decimal point is replaced with a comma to meet SRT specifications (hh:mm:ss,mmm). - Content processing: HTML entities are unescaped using
unescape_html()ontext.firstChild.nodeValue.
srt = ""; seq = 0
for text in texts:
if text.firstChild is None: continue
seq += 1
start = float(text.getAttribute('start'))
dur = float(text.getAttribute('dur')) if text.getAttribute('dur') else 1.0
finish = start + dur
m, s = divmod(start, 60); h, m = divmod(m, 60)
start = '{:0>2}:{:0>2}:{:06.3f}'.format(int(h), int(m), s).replace('.', ',')
m, s = divmod(finish, 60); h, m = divmod(m, 60)
finish = '{:0>2}:{:0>2}:{:06.3f}'.format(int(h), int(m), s).replace('.', ',')
content = unescape_html(text.firstChild.nodeValue)
srt += f'{seq}\n{start} --> {finish}\n{content}\n\n'
The generated SRT string is stored in self.caption_tracks. The extractor distinguishes automatic captions (stored under ct['vssId']) from standard captions (stored under the language code) by checking for the 'kind' key:
if 'kind' in ct:
self.caption_tracks[ct['vssId']] = srt
else:
self.caption_tracks[lang] = srt
Writing SRT Files via the Base Extractor Class
Once video streams complete downloading, the base Extractor class in src/you_get/extractor.py (lines 48-55) handles the generic file-writing logic. This method iterates over all entries in self.caption_tracks and persists them as individual .srt files using UTF-8 encoding.
The filename follows the pattern {sanitized_title}.{language_code}.srt, constructed via get_filename(self.title):
for lang in self.caption_tracks:
filename = f'{get_filename(self.title)}.{lang}.srt'
print(f'Saving {filename} ... ', end='', flush=True)
srt = self.caption_tracks[lang]
with open(os.path.join(kwargs['output_dir'], filename),
'w', encoding='utf-8') as x:
x.write(srt)
print('Done.')
If the user disables captions with --no-caption, this block is skipped entirely.
End-to-End Subtitle Extraction Workflow
The complete process from URL to SRT file follows three distinct phases:
- Extraction Phase: The platform-specific extractor (e.g.,
youtube.py) downloads XML caption data, converts it to SRT format, and populatesself.caption_tracks. - Download Phase: The tool downloads the primary video and audio streams to the output directory.
- Persistence Phase: The base
Extractorclass writes each language track as a separate.srtfile alongside the media file.
Command-Line and Python Usage Examples
Download a YouTube video with English subtitles included:
you-get --caption "https://www.youtube.com/watch?v=VIDEO_ID"
This produces VIDEO_TITLE.en.srt in the output directory.
Programmatically download subtitles using the Python API:
from you_get import common
common.download(
url='https://www.youtube.com/watch?v=VIDEO_ID',
output_dir='./downloads',
caption=True,
merge=True
)
After execution, the SRT file appears in ./downloads with the naming convention Title.Language.srt.
Summary
- XML Parsing: Platform extractors like
src/you_get/extractors/youtube.pyparse caption XML from the video source and calculate precise SRT timestamps usingstartanddurattributes. - SRT Construction: The code formats timestamps as
hh:mm:ss,mmm(comma-separated milliseconds) and unescapes HTML entities to produce valid SubRip content. - File Persistence: The base
Extractorclass insrc/you_get/extractor.pywrites UTF-8 encoded.srtfiles using the pattern{title}.{lang}.srtafter completing video downloads. - Multi-Platform Support: Extractors for YouTube, iQiyi, and Acfun implement the same
self.caption_tracksdictionary interface, ensuring consistent subtitle handling across supported sites.
Frequently Asked Questions
What subtitle format does you-get generate?
you-get generates standard SubRip (SRT) format files. Each file contains sequential entry numbers, time ranges formatted as hh:mm:ss,mmm --> hh:mm:ss,mmm, and subtitle text with HTML entities properly unescaped.
How does you-get calculate SRT timestamps from YouTube captions?
According to the source code in src/you_get/extractors/youtube.py, you-get reads the start and dur (duration) attributes from YouTube's XML <text> elements. It calculates the finish time as start + dur (defaulting to 1 second if duration is missing), then formats both timestamps by replacing the decimal point with a comma to match SRT specifications.
Where does you-get save subtitle files?
Subtitle files are saved in the specified output directory (default is current directory) alongside the video file. The filename follows the pattern {sanitized_video_title}.{language_code}.srt, such as MyDocumentary.en.srt for English captions.
Can you-get download automatic captions from YouTube?
Yes. The YouTube extractor checks for the 'kind' key in caption track metadata. If present, it stores the SRT content under the vssId key (e.g., a.en for automatic English) in self.caption_tracks; otherwise, it uses the standard language code. This allows you-get to handle both human-generated and auto-generated captions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →