Understanding you-get Extractor System Architecture: A Deep Dive into the Plugin Framework
you-get implements a dynamic plugin-style extractor framework that maps URLs to site-specific Python modules inheriting from the abstract VideoExtractor base class, enabling extensible video downloading across hundreds of platforms.
The you-get project is a command-line utility for downloading videos from various websites. Its architecture centers on a flexible extractor system that decouples site-specific parsing logic from core download mechanics, allowing contributors to add support for new platforms without modifying the central dispatcher.
Core Abstract Classes in you-get's Extractor System
The foundation of the architecture rests in src/you_get/extractor.py, which defines the contract that all extractors must follow.
The Extractor Base Class
The minimal Extractor class serves as a generic container for URL, title, and stream information. Defined at lines 10-23 in src/you_get/extractor.py, it provides the basic structure that even non-video extractors can utilize.
The VideoExtractor Workhorse
VideoExtractor is the primary abstract base that powers video downloads. Located in the same file starting at line 21, it implements the template method pattern with three critical hooks that subclasses override:
prepare– Sets up request headers, cookies, and validates the URL.extract– Parses the page or API to populateself.streamsorself.dash_streams.download– Delegates to the common download routine.
The class also provides entry points download_by_url and download_by_vid, along with helper methods p, p_stream, p_i, and p_playlist for formatted output.
class VideoExtractor():
def download_by_url(self, url, **kwargs):
self.url = url
self.prepare(**kwargs) # site-specific preparation
self.streams_sorted = ... # order streams by quality
self.extract(**kwargs) # fill self.streams / self.dash_streams
self.download(**kwargs) # delegate to common download routine
Site-to-Module Mapping and Dynamic Dispatch
The you-get extractor system architecture uses a registry-based approach to route URLs to the correct handler without hard-coding site logic in the main entry point.
The SITES Registry
Located in src/you_get/common.py (lines 25-127), the SITES dictionary maps short host keys like 'youtube', 'bilibili', or 'vimeo' to their corresponding module names under src/you_get/extractors/.
When the CLI invokes common.script_main, the dispatcher extracts the hostname from the input URL, performs a lookup in SITES, and dynamically imports the module using import_module:
module = import_module('.'.join(['you_get', 'extractors', SITES[k]]))
Universal Fallback Extraction
If no exact match exists in the SITES registry, the system falls back to you_get.extractors.universal. This generic extractor attempts to locate video streams through common patterns (direct video file links, embedded players) without site-specific parsing logic.
Concrete Extractor Implementation Pattern
Every site-specific file in src/you_get/extractors/ (such as youtube.py, bilibili.py, or tiktok.py) follows a consistent structural pattern that implements the you-get extractor system architecture.
The prepare() Hook
The prepare method handles pre-flight configuration: setting custom User-Agent headers, initializing cookie jars, extracting video IDs from URL patterns, and handling authentication tokens if required.
The extract() Hook
This is where the site-specific logic resides. The method typically:
- Fetches the page HTML or calls the site's internal API.
- Parses JSON responses or regex-extracts stream data.
- Populates
self.streamswith dictionaries containingsrc(URL list),container(file extension),size, andqualitymetadata.
from ..extractor import VideoExtractor
class YouTube(VideoExtractor):
name = 'youtube'
stream_types = [
{'id': 'flv', 'container': 'flv'},
{'id': 'mp4', 'container': 'mp4'},
]
def prepare(self, **kwargs):
# Set headers, extract video ID from URL
pass
def extract(self, **kwargs):
# Call YouTube API, parse stream maps
# self.streams['1080p'] = {'src': [url1, url2], 'container': 'mp4', 'size': 123456}
pass
Stream Population and Metadata
Concrete extractors populate either self.streams (for progressive downloads) or self.dash_streams (for DASH manifests). Each stream entry must provide the source URLs and container format so that the common download logic in common.py can handle retrieval and file assembly.
Download Flow and Execution Pipeline
The you-get extractor system architecture separates metadata extraction from the actual byte retrieval through a well-defined pipeline:
- CLI Entry:
python -m you_getinvokessrc/you_get/__main__.py, which callscommon.script_main. - URL Dispatch:
script_mainparses arguments, identifies the host, and imports the appropriate extractor module via theSITESregistry. - Extractor Instantiation: The dispatcher calls
download(for single URLs) ordownload_playlist, which instantiates the concrete extractor class. - Template Method Execution:
VideoExtractor.download_by_urlexecutes the sequence:prepare→extract→download. - Common Download Logic: The final
downloadmethod delegates todownload_urlsincommon.py, which handles:- Size estimation via
urls_size - Progress bar creation
- HTTP retrieval via
url_save - Post-processing (FFmpeg merging via
src/you_get/processor/ffmpeg.pyor internal joiners for multi-part downloads)
- Size estimation via
Extending the you-get Extractor Architecture
Adding support for new video platforms requires no modifications to the core dispatcher, demonstrating the plug-in nature of the you-get extractor system architecture:
- Create a new file
src/you_get/extractors/<site>.py. - Define a class inheriting from
VideoExtractor(orExtractorfor non-video content). - Implement the
prepareandextracthooks to populateself.streams. - Add an entry to the
SITESdictionary insrc/you_get/common.pymapping the host key to your module name. - The framework automatically loads and executes the new extractor when matching URLs are processed.
Summary
- you-get employs a plugin-style architecture centered on the
VideoExtractorabstract base class defined insrc/you_get/extractor.py. - The
SITESregistry insrc/you_get/common.pydynamically maps URL hosts to extractor modules usingimport_module. - Concrete extractors implement
prepareandextracthooks to gather metadata and populate stream dictionaries, while the base class handles the standardized download pipeline. - The architecture supports universal fallback extraction and requires no core changes to add new site support, making it highly extensible.
Frequently Asked Questions
How does you-get decide which extractor to use for a given URL?
The dispatcher in src/you_get/common.py extracts the hostname from the input URL and performs a lookup in the SITES dictionary. If a match is found, it dynamically imports the corresponding module from src/you_get/extractors/ using Python's import_module. If no specific match exists, the system falls back to the universal extractor that attempts generic video detection.
What is the difference between the Extractor and VideoExtractor base classes?
Extractor (defined in src/you_get/extractor.py lines 10-23) is a minimal container class used for generic information extraction where video-specific logic is unnecessary. VideoExtractor (starting at line 21 in the same file) extends this functionality with the complete template method pattern including prepare, extract, and download hooks, plus stream management capabilities required for video downloads.
Can I use you-get's extractor system programmatically without the CLI?
Yes, you can instantiate extractor classes directly and invoke their methods. Import the specific extractor from you_get.extractors (e.g., from you_get.extractors.youtube import YouTube), create an instance, and call download_by_url(url, info_only=True) to extract metadata without downloading, or omit info_only to trigger the full download pipeline through the common utilities in src/you_get/common.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →