How to Use youtube-dl Extractors to Add Support for New Video Sites

youtube-dl uses Python extractors that inherit from InfoExtractor to parse video URLs, download metadata, and return standardized format dictionaries, enabling support for any new site by implementing a regex pattern and extraction logic.

The youtube-dl project (ytdl-org/youtube-dl) processes video downloads through a modular extractor architecture. Understanding how to create youtube-dl extractors allows you to extend support to new video hosting platforms by writing minimal Python classes that plug into the existing registry. All extractor logic resides in the youtube_dl/extractor/ directory, with registration handled through a centralized import system.

Architecture of youtube-dl Extractors

Every site supported by youtube-dl is backed by a Python class that knows how to parse that specific platform's URLs and metadata.

The InfoExtractor Base Class

The foundation of every extractor is the InfoExtractor class defined in youtube_dl/extractor/common.py. This base class provides the complete extraction API, including HTTP helper methods (_download_webpage, _download_json), regular expression utilities (_search_regex, _html_search_regex), and error handling. Your concrete extractor must implement two critical attributes: _VALID_URL (a regex pattern matching supported URLs) and _real_extract() (the method that performs the actual extraction).

When youtube-dl runs, the core YoutubeDL class queries list_extractors() (defined in youtube_dl/extractor/__init__.py) to obtain an ordered list of all registered extractor classes. The first extractor whose _VALID_URL pattern matches the input URL wins and executes its _real_extract() method.

Registration and Lazy Loading

The registry in youtube_dl/extractor/extractors.py aggregates import statements for every concrete extractor. To keep startup times fast in recent releases, the project uses a lazy-loading mechanism generated by devscripts/make_lazy_extractors.py, which creates lazy_extractors.py containing a lightweight _ALL_CLASSES list. If lazy loading is active (_LAZY_LOADER is True), you must regenerate this file after adding new extractors.

Step-by-Step Guide to Adding a New Extractor

Follow this checklist to implement support for a new video site. All modifications occur inside the repository working tree.

  1. Create the extractor module. Add a new file named <sitename>.py inside youtube_dl/extractor/.

  2. Import the base class. Add from .common import InfoExtractor at the top of your file (use SearchInfoExtractor instead if the site only provides search functionality).

  3. Define the extractor class. Implement a class with _VALID_URL regex and test cases:

    class MySiteIE(InfoExtractor):
        _VALID_URL = r'https?://(?:www\.)?mysite\.com/watch/(?P<id>\d+)'
        _TESTS = [{
            'url': 'https://www.mysite.com/watch/12345',
            'info_dict': {
                'id': '12345',
                'title': 'example video',
                'ext': 'mp4',
            },
            'params': {'skip_download': True},
        }]
  4. Implement _real_extract(self, url). This method receives the full URL, extracts the video ID using self._match_id(url), and uses helper methods to collect required fields. Return a dictionary containing id, title, and formats following the schema documented in the base class.

  5. Add optional class attributes. Include _AGE_LIMIT for adult content, _GEO_BYPASS or _GEO_COUNTRIES for geo-restricted sites, and IE_DESC for the human-readable description.

  6. Register the extractor. Edit youtube_dl/extractor/extractors.py and add an import line: from .mysite import MySiteIE.

  7. Regenerate lazy extractors. Execute python devscripts/make_lazy_extractors.py youtube_dl/extractor/lazy_extractors.py to update the generated module if lazy loading is enabled.

  8. Write tests. Add entries to the _TESTS list or create test/test_<sitename>.py. The test harness downloads the page (unless skip_download is set) and validates the extracted dictionary.

  9. Update documentation. The IE_DESC class attribute feeds into devscripts/make_readme.py, which generates the supported sites list in the README.

Minimal Working Example

Here is a complete, runnable extractor for a hypothetical example.com site. Save this as youtube_dl/extractor/example.py:

from .common import InfoExtractor

class ExampleSiteIE(InfoExtractor):
    """Extractor for https://example.com/video/<id> pages."""
    IE_DESC = 'Example.com video'
    _VALID_URL = r'https?://(?:www\.)?example\.com/video/(?P<id>\d+)'
    _TESTS = [{
        'url': 'https://example.com/video/42',
        'md5': 'e99a18c428cb38d5f260853678922e03',
        'info_dict': {
            'id': '42',
            'title': 'Demo video',
            'ext': 'mp4',
        },
        'params': {'skip_download': True},
    }]

    def _real_extract(self, url):
        video_id = self._match_id(url)
        webpage = self._download_webpage(url, video_id)

        title = self._html_search_regex(
            r'<h1 class="title">([^<]+)</h1>', webpage, 'title')

        formats = [{
            'url': f'https://media.example.com/video/{video_id}.mp4',
            'ext': 'mp4',
        }]

        return {
            'id': video_id,
            'title': title,
            'formats': formats,
        }

After saving, register the extractor by adding this line to youtube_dl/extractor/extractors.py:

from .example import ExampleSiteIE

Run the test suite to verify functionality:

python -m pytest test/test_example.py

If using the lazy loader, regenerate the extractors list:

python devscripts/make_lazy_extractors.py youtube_dl/extractor/lazy_extractors.py

You can now test the new support with: youtube-dl -F https://example.com/video/42

Essential Source Files for Extractor Development

Understanding these specific files in the ytdl-org/youtube-dl repository is critical for advanced extractor development:

Summary

  • InfoExtractor in youtube_dl/extractor/common.py provides the base API for all extractors, including HTTP utilities and result formatting.
  • New extractors require a _VALID_URL regex pattern and a _real_extract() method that returns a standardized info dictionary.
  • Registration occurs by importing your class in youtube_dl/extractor/extractors.py; no manual list maintenance is required.
  • The lazy loader (make_lazy_extractors.py) must be regenerated after adding extractors to maintain startup performance.
  • Test cases defined in the _TESTS class attribute validate extraction logic against real URLs.

Frequently Asked Questions

What is the minimum code required to create a working youtube-dl extractor?

You need a Python class inheriting from InfoExtractor with two components: a _VALID_URL class attribute containing a regular expression that captures the video ID, and a _real_extract(self, url) method that returns a dictionary with at least id, title, and formats keys. Add the import to extractors.py and the extractor becomes immediately available.

How does youtube-dl decide which extractor to use for a given URL?

The YoutubeDL core calls list_extractors() from youtube_dl/extractor/__init__.py, which returns an ordered list of all registered extractor classes. Each extractor's _VALID_URL pattern is tested against the input URL in sequence, and the first match wins. This is why specific patterns should be listed before generic ones in the extraction order.

Do I need to manually edit the extractor registry when adding a new site?

You only need to add a single import line to youtube_dl/extractor/extractors.py (e.g., from .mysite import MySiteIE). The gen_extractor_classes() function in __init__.py automatically discovers all classes imported there. If the project uses lazy loading, you must also run devscripts/make_lazy_extractors.py to regenerate the compiled extractor list.

What helper methods does InfoExtractor provide for parsing web pages?

The base class offers robust parsing utilities including _download_webpage() for HTTP requests, _download_json() for API endpoints, _search_regex() and _html_search_regex() for pattern matching, _match_id() for extracting the ID from the URL using the _VALID_URL pattern, and _og_search_title() for OpenGraph metadata extraction. These handle retries, encoding, and error reporting automatically.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →