How to Use youtube-dl Extractors to Add Support for New Video Sites
youtube-dl uses Python extractors that inherit from InfoExtractor to parse video URLs, download metadata, and return standardized format dictionaries, enabling support for any new site by implementing a regex pattern and extraction logic.
The youtube-dl project (ytdl-org/youtube-dl) processes video downloads through a modular extractor architecture. Understanding how to create youtube-dl extractors allows you to extend support to new video hosting platforms by writing minimal Python classes that plug into the existing registry. All extractor logic resides in the youtube_dl/extractor/ directory, with registration handled through a centralized import system.
Architecture of youtube-dl Extractors
Every site supported by youtube-dl is backed by a Python class that knows how to parse that specific platform's URLs and metadata.
The InfoExtractor Base Class
The foundation of every extractor is the InfoExtractor class defined in youtube_dl/extractor/common.py. This base class provides the complete extraction API, including HTTP helper methods (_download_webpage, _download_json), regular expression utilities (_search_regex, _html_search_regex), and error handling. Your concrete extractor must implement two critical attributes: _VALID_URL (a regex pattern matching supported URLs) and _real_extract() (the method that performs the actual extraction).
When youtube-dl runs, the core YoutubeDL class queries list_extractors() (defined in youtube_dl/extractor/__init__.py) to obtain an ordered list of all registered extractor classes. The first extractor whose _VALID_URL pattern matches the input URL wins and executes its _real_extract() method.
Registration and Lazy Loading
The registry in youtube_dl/extractor/extractors.py aggregates import statements for every concrete extractor. To keep startup times fast in recent releases, the project uses a lazy-loading mechanism generated by devscripts/make_lazy_extractors.py, which creates lazy_extractors.py containing a lightweight _ALL_CLASSES list. If lazy loading is active (_LAZY_LOADER is True), you must regenerate this file after adding new extractors.
Step-by-Step Guide to Adding a New Extractor
Follow this checklist to implement support for a new video site. All modifications occur inside the repository working tree.
-
Create the extractor module. Add a new file named
<sitename>.pyinsideyoutube_dl/extractor/. -
Import the base class. Add
from .common import InfoExtractorat the top of your file (useSearchInfoExtractorinstead if the site only provides search functionality). -
Define the extractor class. Implement a class with
_VALID_URLregex and test cases:class MySiteIE(InfoExtractor): _VALID_URL = r'https?://(?:www\.)?mysite\.com/watch/(?P<id>\d+)' _TESTS = [{ 'url': 'https://www.mysite.com/watch/12345', 'info_dict': { 'id': '12345', 'title': 'example video', 'ext': 'mp4', }, 'params': {'skip_download': True}, }] -
Implement
_real_extract(self, url). This method receives the full URL, extracts the video ID usingself._match_id(url), and uses helper methods to collect required fields. Return a dictionary containingid,title, andformatsfollowing the schema documented in the base class. -
Add optional class attributes. Include
_AGE_LIMITfor adult content,_GEO_BYPASSor_GEO_COUNTRIESfor geo-restricted sites, andIE_DESCfor the human-readable description. -
Register the extractor. Edit
youtube_dl/extractor/extractors.pyand add an import line:from .mysite import MySiteIE. -
Regenerate lazy extractors. Execute
python devscripts/make_lazy_extractors.py youtube_dl/extractor/lazy_extractors.pyto update the generated module if lazy loading is enabled. -
Write tests. Add entries to the
_TESTSlist or createtest/test_<sitename>.py. The test harness downloads the page (unlessskip_downloadis set) and validates the extracted dictionary. -
Update documentation. The
IE_DESCclass attribute feeds intodevscripts/make_readme.py, which generates the supported sites list in the README.
Minimal Working Example
Here is a complete, runnable extractor for a hypothetical example.com site. Save this as youtube_dl/extractor/example.py:
from .common import InfoExtractor
class ExampleSiteIE(InfoExtractor):
"""Extractor for https://example.com/video/<id> pages."""
IE_DESC = 'Example.com video'
_VALID_URL = r'https?://(?:www\.)?example\.com/video/(?P<id>\d+)'
_TESTS = [{
'url': 'https://example.com/video/42',
'md5': 'e99a18c428cb38d5f260853678922e03',
'info_dict': {
'id': '42',
'title': 'Demo video',
'ext': 'mp4',
},
'params': {'skip_download': True},
}]
def _real_extract(self, url):
video_id = self._match_id(url)
webpage = self._download_webpage(url, video_id)
title = self._html_search_regex(
r'<h1 class="title">([^<]+)</h1>', webpage, 'title')
formats = [{
'url': f'https://media.example.com/video/{video_id}.mp4',
'ext': 'mp4',
}]
return {
'id': video_id,
'title': title,
'formats': formats,
}
After saving, register the extractor by adding this line to youtube_dl/extractor/extractors.py:
from .example import ExampleSiteIE
Run the test suite to verify functionality:
python -m pytest test/test_example.py
If using the lazy loader, regenerate the extractors list:
python devscripts/make_lazy_extractors.py youtube_dl/extractor/lazy_extractors.py
You can now test the new support with: youtube-dl -F https://example.com/video/42
Essential Source Files for Extractor Development
Understanding these specific files in the ytdl-org/youtube-dl repository is critical for advanced extractor development:
youtube_dl/extractor/common.py– Contains theInfoExtractorbase class, HTTP helper methods, and the complete result schema documentation.youtube_dl/extractor/__init__.py– Implementsgen_extractor_classes()andlist_extractors(), building the ordered list of all extractors.youtube_dl/extractor/extractors.py– The aggregation point for all extractor imports; this is where you addfrom .mysite import MySiteIE.devscripts/make_lazy_extractors.py– Generateslazy_extractors.pyto maintain fast import times; must be rerun after adding new extractors when_LAZY_LOADERis active.devscripts/make_readme.py– PullsIE_DESCstrings from all extractors to generate the supported sites documentation.
Summary
- InfoExtractor in
youtube_dl/extractor/common.pyprovides the base API for all extractors, including HTTP utilities and result formatting. - New extractors require a
_VALID_URLregex pattern and a_real_extract()method that returns a standardized info dictionary. - Registration occurs by importing your class in
youtube_dl/extractor/extractors.py; no manual list maintenance is required. - The lazy loader (
make_lazy_extractors.py) must be regenerated after adding extractors to maintain startup performance. - Test cases defined in the
_TESTSclass attribute validate extraction logic against real URLs.
Frequently Asked Questions
What is the minimum code required to create a working youtube-dl extractor?
You need a Python class inheriting from InfoExtractor with two components: a _VALID_URL class attribute containing a regular expression that captures the video ID, and a _real_extract(self, url) method that returns a dictionary with at least id, title, and formats keys. Add the import to extractors.py and the extractor becomes immediately available.
How does youtube-dl decide which extractor to use for a given URL?
The YoutubeDL core calls list_extractors() from youtube_dl/extractor/__init__.py, which returns an ordered list of all registered extractor classes. Each extractor's _VALID_URL pattern is tested against the input URL in sequence, and the first match wins. This is why specific patterns should be listed before generic ones in the extraction order.
Do I need to manually edit the extractor registry when adding a new site?
You only need to add a single import line to youtube_dl/extractor/extractors.py (e.g., from .mysite import MySiteIE). The gen_extractor_classes() function in __init__.py automatically discovers all classes imported there. If the project uses lazy loading, you must also run devscripts/make_lazy_extractors.py to regenerate the compiled extractor list.
What helper methods does InfoExtractor provide for parsing web pages?
The base class offers robust parsing utilities including _download_webpage() for HTTP requests, _download_json() for API endpoints, _search_regex() and _html_search_regex() for pattern matching, _match_id() for extracting the ID from the URL using the _VALID_URL pattern, and _og_search_title() for OpenGraph metadata extraction. These handle retries, encoding, and error reporting automatically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →