Understanding you-get Extractor System Architecture: A Deep Dive into the Plugin Framework

you-get implements a dynamic plugin-style extractor framework that maps URLs to site-specific Python modules inheriting from the abstract VideoExtractor base class, enabling extensible video downloading across hundreds of platforms.

The you-get project is a command-line utility for downloading videos from various websites. Its architecture centers on a flexible extractor system that decouples site-specific parsing logic from core download mechanics, allowing contributors to add support for new platforms without modifying the central dispatcher.

Core Abstract Classes in you-get's Extractor System

The foundation of the architecture rests in src/you_get/extractor.py, which defines the contract that all extractors must follow.

The Extractor Base Class

The minimal Extractor class serves as a generic container for URL, title, and stream information. Defined at lines 10-23 in src/you_get/extractor.py, it provides the basic structure that even non-video extractors can utilize.

The VideoExtractor Workhorse

VideoExtractor is the primary abstract base that powers video downloads. Located in the same file starting at line 21, it implements the template method pattern with three critical hooks that subclasses override:

  • prepare – Sets up request headers, cookies, and validates the URL.
  • extract – Parses the page or API to populate self.streams or self.dash_streams.
  • download – Delegates to the common download routine.

The class also provides entry points download_by_url and download_by_vid, along with helper methods p, p_stream, p_i, and p_playlist for formatted output.

class VideoExtractor():
    def download_by_url(self, url, **kwargs):
        self.url = url
        self.prepare(**kwargs)          # site-specific preparation

        self.streams_sorted = ...         # order streams by quality

        self.extract(**kwargs)           # fill self.streams / self.dash_streams

        self.download(**kwargs)          # delegate to common download routine

Site-to-Module Mapping and Dynamic Dispatch

The you-get extractor system architecture uses a registry-based approach to route URLs to the correct handler without hard-coding site logic in the main entry point.

The SITES Registry

Located in src/you_get/common.py (lines 25-127), the SITES dictionary maps short host keys like 'youtube', 'bilibili', or 'vimeo' to their corresponding module names under src/you_get/extractors/.

When the CLI invokes common.script_main, the dispatcher extracts the hostname from the input URL, performs a lookup in SITES, and dynamically imports the module using import_module:

module = import_module('.'.join(['you_get', 'extractors', SITES[k]]))

Universal Fallback Extraction

If no exact match exists in the SITES registry, the system falls back to you_get.extractors.universal. This generic extractor attempts to locate video streams through common patterns (direct video file links, embedded players) without site-specific parsing logic.

Concrete Extractor Implementation Pattern

Every site-specific file in src/you_get/extractors/ (such as youtube.py, bilibili.py, or tiktok.py) follows a consistent structural pattern that implements the you-get extractor system architecture.

The prepare() Hook

The prepare method handles pre-flight configuration: setting custom User-Agent headers, initializing cookie jars, extracting video IDs from URL patterns, and handling authentication tokens if required.

The extract() Hook

This is where the site-specific logic resides. The method typically:

  1. Fetches the page HTML or calls the site's internal API.
  2. Parses JSON responses or regex-extracts stream data.
  3. Populates self.streams with dictionaries containing src (URL list), container (file extension), size, and quality metadata.
from ..extractor import VideoExtractor

class YouTube(VideoExtractor):
    name = 'youtube'
    stream_types = [
        {'id': 'flv', 'container': 'flv'},
        {'id': 'mp4', 'container': 'mp4'},
    ]

    def prepare(self, **kwargs):
        # Set headers, extract video ID from URL

        pass

    def extract(self, **kwargs):
        # Call YouTube API, parse stream maps

        # self.streams['1080p'] = {'src': [url1, url2], 'container': 'mp4', 'size': 123456}

        pass

Stream Population and Metadata

Concrete extractors populate either self.streams (for progressive downloads) or self.dash_streams (for DASH manifests). Each stream entry must provide the source URLs and container format so that the common download logic in common.py can handle retrieval and file assembly.

Download Flow and Execution Pipeline

The you-get extractor system architecture separates metadata extraction from the actual byte retrieval through a well-defined pipeline:

  1. CLI Entry: python -m you_get invokes src/you_get/__main__.py, which calls common.script_main.
  2. URL Dispatch: script_main parses arguments, identifies the host, and imports the appropriate extractor module via the SITES registry.
  3. Extractor Instantiation: The dispatcher calls download (for single URLs) or download_playlist, which instantiates the concrete extractor class.
  4. Template Method Execution: VideoExtractor.download_by_url executes the sequence: prepare → extract → download.
  5. Common Download Logic: The final download method delegates to download_urls in common.py, which handles:
    • Size estimation via urls_size
    • Progress bar creation
    • HTTP retrieval via url_save
    • Post-processing (FFmpeg merging via src/you_get/processor/ffmpeg.py or internal joiners for multi-part downloads)

Extending the you-get Extractor Architecture

Adding support for new video platforms requires no modifications to the core dispatcher, demonstrating the plug-in nature of the you-get extractor system architecture:

  1. Create a new file src/you_get/extractors/<site>.py.
  2. Define a class inheriting from VideoExtractor (or Extractor for non-video content).
  3. Implement the prepare and extract hooks to populate self.streams.
  4. Add an entry to the SITES dictionary in src/you_get/common.py mapping the host key to your module name.
  5. The framework automatically loads and executes the new extractor when matching URLs are processed.

Summary

  • you-get employs a plugin-style architecture centered on the VideoExtractor abstract base class defined in src/you_get/extractor.py.
  • The SITES registry in src/you_get/common.py dynamically maps URL hosts to extractor modules using import_module.
  • Concrete extractors implement prepare and extract hooks to gather metadata and populate stream dictionaries, while the base class handles the standardized download pipeline.
  • The architecture supports universal fallback extraction and requires no core changes to add new site support, making it highly extensible.

Frequently Asked Questions

How does you-get decide which extractor to use for a given URL?

The dispatcher in src/you_get/common.py extracts the hostname from the input URL and performs a lookup in the SITES dictionary. If a match is found, it dynamically imports the corresponding module from src/you_get/extractors/ using Python's import_module. If no specific match exists, the system falls back to the universal extractor that attempts generic video detection.

What is the difference between the Extractor and VideoExtractor base classes?

Extractor (defined in src/you_get/extractor.py lines 10-23) is a minimal container class used for generic information extraction where video-specific logic is unnecessary. VideoExtractor (starting at line 21 in the same file) extends this functionality with the complete template method pattern including prepare, extract, and download hooks, plus stream management capabilities required for video downloads.

Can I use you-get's extractor system programmatically without the CLI?

Yes, you can instantiate extractor classes directly and invoke their methods. Import the specific extractor from you_get.extractors (e.g., from you_get.extractors.youtube import YouTube), create an instance, and call download_by_url(url, info_only=True) to extract metadata without downloading, or omit info_only to trigger the full download pipeline through the common utilities in src/you_get/common.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →