# How you-get Detectes Sites and Matches URL Patterns Internally

> Discover how you-get detects sites and matches URL patterns internally. Learn about its two-stage pipeline domain mapping and regex validation for efficient video downloading.

- Repository: [Mort Yao/you-get](https://github.com/soimort/you-get)
- Tags: internals
- Published: 2026-03-06

---

**you-get uses a two-stage pipeline that maps domains to extractor modules via a `SITES` dictionary, then validates URLs using regex patterns through a `match1` helper function.**

Understanding how you-get detects sites and matches URL patterns internally is essential for developers who want to extend the tool or debug extraction failures. The popular Python video downloader delegates site-specific logic to dedicated extractor modules, but first it must determine which module should handle a given URL. This detection happens through a lightweight yet extensible routing system implemented in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) and the individual extractor classes.

## The Two-Stage Detection Pipeline

The internal architecture separates site detection from URL validation. This separation allows you-get to quickly narrow down candidate extractors using simple string matching, then apply rigorous regex validation only on the relevant module.

### Stage 1: Domain-to-Extractor Mapping via SITES

At the heart of the routing system lies the **`SITES`** dictionary defined in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) (lines 25-27). This mapping uses short identifier strings as keys and module names as values:

```python
SITES = {
    'youtube': 'youtube',
    'zhihu': 'zhihu',
    'bilibili': 'bilibili',
    # ... additional entries

}

```

When you-get receives a URL, it iterates over this dictionary and performs a case-insensitive substring search. If the key appears anywhere in the URL's hostname or path, you-get knows which extractor module to import.

### Stage 2: Pattern Validation with match1

After identifying the candidate module, you-get instantiates the extractor class and delegates final URL validation to it. Each extractor inherits from the base `Extractor` class defined in [`src/you_get/extractor.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractor.py) and implements site-specific logic using the **`match1`** helper function.

## Deep Dive into the SITES Dictionary

The `SITES` dictionary acts as a lightweight router that avoids expensive regex operations during the initial screening phase. Located in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py), this data structure maps domain fragments to their corresponding Python modules in `src/you_get/extractors/`.

```python
from importlib import import_module  # line 15 in common.py

SITES = {
    'youtube': 'youtube',
    'youku': 'youku',
    'tudou': 'tudou',
    # Additional mappings...

}

```

The dispatcher logic iterates over these entries using a simple membership test:

```python
for key, module_name in SITES.items():
    if key in url.lower():
        mod = import_module(f'you_get.extractors.{module_name}')
        extractor = mod.Extractor(url)
        extractor.download_by_url(url, **kwargs)
        break
else:
    # Fallback to universal extractor

    from you_get.extractors.universal import universal_download
    universal_download(url, **kwargs)

```

This approach ensures that only one extractor module is loaded per URL, keeping memory usage minimal and startup times fast.

## How Extractors Validate URLs

Once the dispatcher identifies a candidate module, the actual URL pattern matching occurs within the extractor class itself. Each site-specific extractor defines regex patterns that validate the URL structure and extract essential identifiers like video IDs.

### The match1 Helper Function

The `match1` function, defined in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) (lines 26-38), serves as the primary regex utility across all extractors. It accepts an input string and multiple regex patterns, returning the first captured group that matches:

```python
import re

def match1(text, *patterns):
    """Scans through a string for any of multiple patterns."""
    for pattern in patterns:
        match = re.search(pattern, text)
        if match:
            return match.group(1)
    return None

```

This function enables extractors to test multiple URL formats efficiently without chaining complex conditional statements.

### Pattern Matching in Practice

Consider the YouTube extractor in [`src/you_get/extractors/youtube.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractors/youtube.py). It uses `match1` to recognize various YouTube URL formats and extract the video identifier:

```python
class Youtube(Extractor):
    @classmethod
    def get_vid_from_url(cls, url):
        """Extract video ID from various YouTube URL patterns."""
        return match1(
            url,
            r'youtu\.be/([^?/]+)',                     # Short URL format

            r'youtube\.com/embed/([^/?]+)',            # Embedded player

            r'youtube\.com/shorts/([^/?]+)',           # YouTube Shorts

            r'youtube\.com/watch\?v=([^&]+)',          # Standard watch URL

            parse_query_param(url, 'v')                # Fallback to query parser

        )

```

The first pattern that successfully captures a video ID determines the URL type, allowing the extractor to proceed with format-specific logic. If `match1` returns `None`, the extractor typically raises an exception indicating that the URL is unsupported.

## Dynamic Module Loading and Execution

The final piece of the detection pipeline involves dynamically importing the appropriate extractor module and instantiating its class. This occurs in the main download dispatcher within [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py).

When the `SITES` lookup identifies a matching key, you-get uses Python's `importlib.import_module` to load the extractor on demand:

```python
from importlib import import_module

def download_main(url, **kwargs):
    # ... initialization code ...

    
    for key, module_name in SITES.items():
        if key in url.lower():
            # Dynamic import: you_get.extractors.youtube, etc.

            mod = import_module(f'you_get.extractors.{module_name}')
            
            # Instantiate the specific extractor class

            extractor = mod.Extractor(url)
            
            # Delegate to the extractor's download method

            extractor.download_by_url(url, **kwargs)
            return True
    
    # No specific extractor found; use universal fallback

    from you_get.extractors.universal import universal_download
    universal_download(url, **kwargs)
    return True

```

This lazy-loading approach ensures that you-get only imports the code necessary for the specific site being processed, reducing startup overhead and memory consumption. If no pattern in `SITES` matches the input URL, the system falls back to the `universal` extractor, which attempts generic media discovery through page scraping.

## Summary

- **you-get** routes URLs using a two-stage detection system that balances speed with accuracy.
- The **`SITES`** dictionary in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) provides coarse-grained domain matching to identify candidate extractor modules.
- The **`match1`** helper function enables fine-grained regex validation within individual extractors, supporting multiple URL formats per site.
- Extractor modules are loaded dynamically using `importlib.import_module` only when needed, optimizing performance.
- If no specific extractor matches, you-get falls back to the `universal` extractor for generic media discovery.

## Frequently Asked Questions

### How does you-get decide which extractor to use for a URL?

you-get first checks the **`SITES`** dictionary in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) to find a domain keyword that appears in the URL. If a match is found, it dynamically imports the corresponding extractor module from `src/you_get/extractors/`. The extractor's internal regex patterns then validate that the specific URL format is supported.

### What happens if you-get doesn't recognize a website?

If no key in the `SITES` dictionary matches the input URL, you-get falls back to the **`universal`** extractor defined in [`src/you_get/extractors/universal.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractors/universal.py). This generic extractor attempts to discover media URLs by parsing the page's HTML and searching for common video file patterns, though it may fail on sites with complex authentication or encryption.

### Can I add support for a new website by modifying the SITES dictionary?

Yes, but you must also create a corresponding extractor module. Adding an entry to the `SITES` dictionary in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) (e.g., `'newsite': 'newsite'`) tells you-get to import [`src/you_get/extractors/newsite.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractors/newsite.py). That module must define an `Extractor` class that implements the `download_by_url` method and uses `match1` to validate URL patterns specific to that site.

### What is the match1 function used for in you-get?

The **`match1`** function in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) is a utility that applies multiple regular expression patterns to a string and returns the first captured group that matches. Extractors use it to extract video IDs from various URL formats (such as short links, embed URLs, and standard watch pages) without writing complex conditional logic. It returns `None` if no patterns match, signaling that the URL format is unsupported.