Site-Specific Extractors in you-get: Structure and Interface Implementation

Site-specific extractors in you-get are Python classes that inherit from VideoExtractor and implement the extractor interface by overriding prepare() and extract() methods to parse site-specific media URLs.

The you-get tool organizes video extraction logic into modular components located in src/you_get/extractors/. Each site-specific extractor implements a standardized interface defined in src/you_get/extractor.py, enabling the framework to download content from hundreds of supported sites through a unified architecture.

The Extractor Interface and Base Class

All site-specific extractors inherit from the abstract VideoExtractor base class defined in src/you_get/extractor.py. This base class provides the Extractor interface—a contract that every concrete extractor must fulfill by implementing specific methods.

The base class supplies these core methods:

Method Purpose
download_by_url(url, **kwargs) Entry point called by the CLI. Sets self.url, runs prepare(), builds self.streams_sorted, calls extract(), then download().
download_by_vid(vid, **kwargs) Alternative entry point that accepts a video ID instead of a full URL.
prepare(**kwargs) Must be overridden—fetches the page, extracts the title, and populates self.streams with raw URLs.
extract(**kwargs) Must be overridden—transforms raw entries in self.streams into fully-described streams with container, size, and source lists.
download(**kwargs) Shared implementation handling JSON output, info_only mode, and actual downloading.

Minimal Extractor Structure

A basic site-specific extractor requires only a class name, a name attribute, a stream_types list, and implementations of prepare() and extract(). The following excerpt from src/you_get/extractors/pinterest.py demonstrates the minimal viable implementation:

from ..common import *
from ..extractor import VideoExtractor

class Pinterest(VideoExtractor):
    name = "Pinterest"

    stream_types = [
        {'id': 'original'},
        {'id': 'small'},
    ]

    def prepare(self, **kwargs):
        content = get_content(self.url)

        self.title = match1(content,
                           r'<meta property="og:description" '
                           r'name="og:description" content="([^"]+)"')

        orig_img = match1(content,
                          r'<meta itemprop="image" content="([^"]+/originals/[^"]+)"')
        twit_img = match1(content,
                          r'<meta property="twitter:image:src" '
                          r'name="twitter:image:src" content="([^"]+)"')

        if orig_img: self.streams['original'] = {'url': orig_img}
        if twit_img: self.streams['small']    = {'url': twit_img}

    def extract(self, **kwargs):
        for i in self.streams:
            s = self.streams[i]
            _, s['container'], s['size'] = url_info(s['url'])
            s['src'] = [s['url']]

This extractor implements the interface solely through prepare and extract. All orchestration logic, including stream selection and download management, is inherited from VideoExtractor.

Complex Site Implementations

For sites requiring signed URLs, API calls, or signature decoding (such as YouTube or Bilibili), site-specific extractors extend the basic pattern with helper methods and detailed stream_types configurations. These extractors still override only prepare() and extract(), but add internal methods to handle site-specific encryption or throttling.

Key attributes for complex extractors include:

  • stream_types—A list of dictionaries containing itag or id, container, video_resolution, and quality keys that define quality preference order.
  • Static helpers—Methods that build API URLs or decode signatures (e.g., s_to_sig, dethrottle in the YouTube extractor).
  • Custom prepare logic—Parses JSON embedded in pages, makes additional HTTP requests, or handles authentication flows.

The YouTube extractor in src/you_get/extractors/youtube.py demonstrates this pattern:

class YouTube(VideoExtractor):
    name = "YouTube"

    stream_types = [
        {'itag': '38', 'container': 'MP4', 'video_resolution': '3072p'},
        {'itag': '22', 'container': 'MP4', 'video_resolution': '720p'},
    ]

    def dethrottle(js, url):
        # Rewrites throttled URLs by manipulating the "n" parameter

        return u._replace(query=urlencode(qs, doseq=True)).geturl()

    def s_to_sig(js, s):
        # Converts "s" cipher parameter to "sig" value

        return sig

Despite additional complexity, the core interface remains unchanged: prepare populates self.title and self.streams, while extract enriches stream metadata.

Dynamic Extractor Discovery

The framework discovers available site-specific extractors dynamically through src/you_get/common.py. The system walks the src/you_get/extractors/ package, imports every module, and registers classes that expose a name attribute matching supported site domains.

The registration logic functions approximately as follows:

def get_extractor(url):
    host = parse_host(url)
    for mod in pkgutil.iter_modules(extractors.__path__):
        module = import_module('you_get.extractors.' + mod.name)
        for name, cls in inspect.getmembers(module, inspect.isclass):
            if getattr(cls, 'name', None) and host in cls.sites:
                return cls()
    raise NotImplementedError(...)

Adding a new site requires only placing a new subclass in src/you_get/extractors/ that inherits from VideoExtractor and exposes the name attribute—the framework handles registration automatically.

Practical Usage Examples

All site-specific extractors expose the same public interface regardless of internal complexity:

>>> from you_get.extractors.youtube import YouTube
>>> yt = YouTube()
>>> yt.download_by_url('https://www.youtube.com/watch?v=dQw4w9WgXcQ',
...                    info_only=True)
site:                YouTube
title:               Rick Astley - Never Gonna Give You Up (Official Music Video)
stream:              # Best quality

    - itag:          22
      container:     MP4
      quality:       720p
      size:          15.2 MiB (15972320 bytes)
>>> from you_get.extractors.pinterest import Pinterest
>>> pin = Pinterest()
>>> pin.download_by_url('https://www.pinterest.com/pin/123456789012345678/',
...                    info_only=True)
site:                Pinterest
title:               Cute cat picture
stream:              # Best quality

    - format:        original
      container:     jpg
      size:          0.8 MiB (839000 bytes)

Both examples utilize download_by_url() identically, even though the underlying implementations differ significantly.

Key Files in the Extractor Architecture

File Role
src/you_get/extractor.py Defines Extractor and VideoExtractor base classes establishing the interface contract.
src/you_get/common.py Provides utility functions (get_content, match1, url_info) and the dynamic extractor loader.
src/you_get/extractors/*.py Site-specific implementations; each contains a class inheriting from VideoExtractor with name, stream_types, and overridden methods.
src/you_get/util/*.py Low-level helpers for HTTP handling, logging, and JSON output used by extractors.

Summary

  • Site-specific extractors are subclasses of VideoExtractor located in src/you_get/extractors/.
  • Mandatory implementation requires overriding prepare() to fetch metadata and extract() to populate stream details.
  • Standard attributes include name (site identifier) and stream_types (quality ordering).
  • Dynamic discovery occurs automatically via src/you_get/common.py by inspecting the extractors package.
  • Complex sites add internal helper methods while maintaining the same public interface as simple extractors.

Frequently Asked Questions

What methods must a site-specific extractor override?

A site-specific extractor must override prepare() to fetch the page and populate self.title and self.streams, and extract() to enrich each stream with container format, file size, and source URLs. All other methods, including download_by_url() and download(), inherit their behavior from VideoExtractor.

How does you-get map URLs to extractors?

The framework uses dynamic discovery in src/you_get/common.py to walk the src/you_get/extractors/ directory, import each module, and match the URL host against the sites attribute of each extractor class. When matched, it instantiates the corresponding class automatically.

What is the purpose of the stream_types attribute?

The stream_types attribute defines an ordered list of supported quality identifiers for a site. Each dictionary must contain either id or itag keys, and may specify container, video_resolution, and quality. The base class uses this list to sort available streams from highest to lowest quality.

Can extractors implement additional helper methods?

Yes, extractors frequently implement private helper methods for site-specific logic such as signature decoding, API URL construction, or throttling removal. These helpers are internal implementation details and do not change the public interface, which remains consistent across all extractors through download_by_url().

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →