How you-get Detectes Sites and Matches URL Patterns Internally
you-get uses a two-stage pipeline that maps domains to extractor modules via a SITES dictionary, then validates URLs using regex patterns through a match1 helper function.
Understanding how you-get detects sites and matches URL patterns internally is essential for developers who want to extend the tool or debug extraction failures. The popular Python video downloader delegates site-specific logic to dedicated extractor modules, but first it must determine which module should handle a given URL. This detection happens through a lightweight yet extensible routing system implemented in src/you_get/common.py and the individual extractor classes.
The Two-Stage Detection Pipeline
The internal architecture separates site detection from URL validation. This separation allows you-get to quickly narrow down candidate extractors using simple string matching, then apply rigorous regex validation only on the relevant module.
Stage 1: Domain-to-Extractor Mapping via SITES
At the heart of the routing system lies the SITES dictionary defined in src/you_get/common.py (lines 25-27). This mapping uses short identifier strings as keys and module names as values:
SITES = {
'youtube': 'youtube',
'zhihu': 'zhihu',
'bilibili': 'bilibili',
# ... additional entries
}
When you-get receives a URL, it iterates over this dictionary and performs a case-insensitive substring search. If the key appears anywhere in the URL's hostname or path, you-get knows which extractor module to import.
Stage 2: Pattern Validation with match1
After identifying the candidate module, you-get instantiates the extractor class and delegates final URL validation to it. Each extractor inherits from the base Extractor class defined in src/you_get/extractor.py and implements site-specific logic using the match1 helper function.
Deep Dive into the SITES Dictionary
The SITES dictionary acts as a lightweight router that avoids expensive regex operations during the initial screening phase. Located in src/you_get/common.py, this data structure maps domain fragments to their corresponding Python modules in src/you_get/extractors/.
from importlib import import_module # line 15 in common.py
SITES = {
'youtube': 'youtube',
'youku': 'youku',
'tudou': 'tudou',
# Additional mappings...
}
The dispatcher logic iterates over these entries using a simple membership test:
for key, module_name in SITES.items():
if key in url.lower():
mod = import_module(f'you_get.extractors.{module_name}')
extractor = mod.Extractor(url)
extractor.download_by_url(url, **kwargs)
break
else:
# Fallback to universal extractor
from you_get.extractors.universal import universal_download
universal_download(url, **kwargs)
This approach ensures that only one extractor module is loaded per URL, keeping memory usage minimal and startup times fast.
How Extractors Validate URLs
Once the dispatcher identifies a candidate module, the actual URL pattern matching occurs within the extractor class itself. Each site-specific extractor defines regex patterns that validate the URL structure and extract essential identifiers like video IDs.
The match1 Helper Function
The match1 function, defined in src/you_get/common.py (lines 26-38), serves as the primary regex utility across all extractors. It accepts an input string and multiple regex patterns, returning the first captured group that matches:
import re
def match1(text, *patterns):
"""Scans through a string for any of multiple patterns."""
for pattern in patterns:
match = re.search(pattern, text)
if match:
return match.group(1)
return None
This function enables extractors to test multiple URL formats efficiently without chaining complex conditional statements.
Pattern Matching in Practice
Consider the YouTube extractor in src/you_get/extractors/youtube.py. It uses match1 to recognize various YouTube URL formats and extract the video identifier:
class Youtube(Extractor):
@classmethod
def get_vid_from_url(cls, url):
"""Extract video ID from various YouTube URL patterns."""
return match1(
url,
r'youtu\.be/([^?/]+)', # Short URL format
r'youtube\.com/embed/([^/?]+)', # Embedded player
r'youtube\.com/shorts/([^/?]+)', # YouTube Shorts
r'youtube\.com/watch\?v=([^&]+)', # Standard watch URL
parse_query_param(url, 'v') # Fallback to query parser
)
The first pattern that successfully captures a video ID determines the URL type, allowing the extractor to proceed with format-specific logic. If match1 returns None, the extractor typically raises an exception indicating that the URL is unsupported.
Dynamic Module Loading and Execution
The final piece of the detection pipeline involves dynamically importing the appropriate extractor module and instantiating its class. This occurs in the main download dispatcher within src/you_get/common.py.
When the SITES lookup identifies a matching key, you-get uses Python's importlib.import_module to load the extractor on demand:
from importlib import import_module
def download_main(url, **kwargs):
# ... initialization code ...
for key, module_name in SITES.items():
if key in url.lower():
# Dynamic import: you_get.extractors.youtube, etc.
mod = import_module(f'you_get.extractors.{module_name}')
# Instantiate the specific extractor class
extractor = mod.Extractor(url)
# Delegate to the extractor's download method
extractor.download_by_url(url, **kwargs)
return True
# No specific extractor found; use universal fallback
from you_get.extractors.universal import universal_download
universal_download(url, **kwargs)
return True
This lazy-loading approach ensures that you-get only imports the code necessary for the specific site being processed, reducing startup overhead and memory consumption. If no pattern in SITES matches the input URL, the system falls back to the universal extractor, which attempts generic media discovery through page scraping.
Summary
- you-get routes URLs using a two-stage detection system that balances speed with accuracy.
- The
SITESdictionary insrc/you_get/common.pyprovides coarse-grained domain matching to identify candidate extractor modules. - The
match1helper function enables fine-grained regex validation within individual extractors, supporting multiple URL formats per site. - Extractor modules are loaded dynamically using
importlib.import_moduleonly when needed, optimizing performance. - If no specific extractor matches, you-get falls back to the
universalextractor for generic media discovery.
Frequently Asked Questions
How does you-get decide which extractor to use for a URL?
you-get first checks the SITES dictionary in src/you_get/common.py to find a domain keyword that appears in the URL. If a match is found, it dynamically imports the corresponding extractor module from src/you_get/extractors/. The extractor's internal regex patterns then validate that the specific URL format is supported.
What happens if you-get doesn't recognize a website?
If no key in the SITES dictionary matches the input URL, you-get falls back to the universal extractor defined in src/you_get/extractors/universal.py. This generic extractor attempts to discover media URLs by parsing the page's HTML and searching for common video file patterns, though it may fail on sites with complex authentication or encryption.
Can I add support for a new website by modifying the SITES dictionary?
Yes, but you must also create a corresponding extractor module. Adding an entry to the SITES dictionary in src/you_get/common.py (e.g., 'newsite': 'newsite') tells you-get to import src/you_get/extractors/newsite.py. That module must define an Extractor class that implements the download_by_url method and uses match1 to validate URL patterns specific to that site.
What is the match1 function used for in you-get?
The match1 function in src/you_get/common.py is a utility that applies multiple regular expression patterns to a string and returns the first captured group that matches. Extractors use it to extract video IDs from various URL formats (such as short links, embed URLs, and standard watch pages) without writing complex conditional logic. It returns None if no patterns match, signaling that the URL format is unsupported.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →