# Site-Specific Extractors in you-get: Structure and Interface Implementation

> Learn the structure of site-specific extractors in you-get and how they implement the VideoExtractor interface by overriding prepare and extract methods.

- Repository: [Mort Yao/you-get](https://github.com/soimort/you-get)
- Tags: internals
- Published: 2026-03-06

---

**Site-specific extractors in you-get are Python classes that inherit from `VideoExtractor` and implement the extractor interface by overriding `prepare()` and `extract()` methods to parse site-specific media URLs.**

The `you-get` tool organizes video extraction logic into modular components located in `src/you_get/extractors/`. Each site-specific extractor implements a standardized interface defined in [`src/you_get/extractor.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractor.py), enabling the framework to download content from hundreds of supported sites through a unified architecture.

## The Extractor Interface and Base Class

All site-specific extractors inherit from the abstract `VideoExtractor` base class defined in [`src/you_get/extractor.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractor.py). This base class provides the **Extractor interface**—a contract that every concrete extractor must fulfill by implementing specific methods.

The base class supplies these core methods:

| Method | Purpose |
|--------|---------|
| `download_by_url(url, **kwargs)` | Entry point called by the CLI. Sets `self.url`, runs `prepare()`, builds `self.streams_sorted`, calls `extract()`, then `download()`. |
| `download_by_vid(vid, **kwargs)` | Alternative entry point that accepts a video ID instead of a full URL. |
| `prepare(**kwargs)` | **Must be overridden**—fetches the page, extracts the title, and populates `self.streams` with raw URLs. |
| `extract(**kwargs)` | **Must be overridden**—transforms raw entries in `self.streams` into fully-described streams with container, size, and source lists. |
| `download(**kwargs)` | Shared implementation handling JSON output, `info_only` mode, and actual downloading. |

## Minimal Extractor Structure

A basic site-specific extractor requires only a class name, a `name` attribute, a `stream_types` list, and implementations of `prepare()` and `extract()`. The following excerpt from [`src/you_get/extractors/pinterest.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractors/pinterest.py) demonstrates the minimal viable implementation:

```python
from ..common import *
from ..extractor import VideoExtractor

class Pinterest(VideoExtractor):
    name = "Pinterest"

    stream_types = [
        {'id': 'original'},
        {'id': 'small'},
    ]

    def prepare(self, **kwargs):
        content = get_content(self.url)

        self.title = match1(content,
                           r'<meta property="og:description" '
                           r'name="og:description" content="([^"]+)"')

        orig_img = match1(content,
                          r'<meta itemprop="image" content="([^"]+/originals/[^"]+)"')
        twit_img = match1(content,
                          r'<meta property="twitter:image:src" '
                          r'name="twitter:image:src" content="([^"]+)"')

        if orig_img: self.streams['original'] = {'url': orig_img}
        if twit_img: self.streams['small']    = {'url': twit_img}

    def extract(self, **kwargs):
        for i in self.streams:
            s = self.streams[i]
            _, s['container'], s['size'] = url_info(s['url'])
            s['src'] = [s['url']]

```

This extractor implements the interface solely through `prepare` and `extract`. All orchestration logic, including stream selection and download management, is inherited from `VideoExtractor`.

## Complex Site Implementations

For sites requiring signed URLs, API calls, or signature decoding (such as YouTube or Bilibili), site-specific extractors extend the basic pattern with helper methods and detailed `stream_types` configurations. These extractors still override only `prepare()` and `extract()`, but add internal methods to handle site-specific encryption or throttling.

Key attributes for complex extractors include:

- **`stream_types`**—A list of dictionaries containing `itag` or `id`, `container`, `video_resolution`, and `quality` keys that define quality preference order.
- **Static helpers**—Methods that build API URLs or decode signatures (e.g., `s_to_sig`, `dethrottle` in the YouTube extractor).
- **Custom `prepare` logic**—Parses JSON embedded in pages, makes additional HTTP requests, or handles authentication flows.

The YouTube extractor in [`src/you_get/extractors/youtube.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractors/youtube.py) demonstrates this pattern:

```python
class YouTube(VideoExtractor):
    name = "YouTube"

    stream_types = [
        {'itag': '38', 'container': 'MP4', 'video_resolution': '3072p'},
        {'itag': '22', 'container': 'MP4', 'video_resolution': '720p'},
    ]

    def dethrottle(js, url):
        # Rewrites throttled URLs by manipulating the "n" parameter

        return u._replace(query=urlencode(qs, doseq=True)).geturl()

    def s_to_sig(js, s):
        # Converts "s" cipher parameter to "sig" value

        return sig

```

Despite additional complexity, the **core interface remains unchanged**: `prepare` populates `self.title` and `self.streams`, while `extract` enriches stream metadata.

## Dynamic Extractor Discovery

The framework discovers available site-specific extractors dynamically through [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py). The system walks the `src/you_get/extractors/` package, imports every module, and registers classes that expose a `name` attribute matching supported site domains.

The registration logic functions approximately as follows:

```python
def get_extractor(url):
    host = parse_host(url)
    for mod in pkgutil.iter_modules(extractors.__path__):
        module = import_module('you_get.extractors.' + mod.name)
        for name, cls in inspect.getmembers(module, inspect.isclass):
            if getattr(cls, 'name', None) and host in cls.sites:
                return cls()
    raise NotImplementedError(...)

```

**Adding a new site requires only placing a new subclass in `src/you_get/extractors/`** that inherits from `VideoExtractor` and exposes the `name` attribute—the framework handles registration automatically.

## Practical Usage Examples

All site-specific extractors expose the same public interface regardless of internal complexity:

```python
>>> from you_get.extractors.youtube import YouTube
>>> yt = YouTube()
>>> yt.download_by_url('https://www.youtube.com/watch?v=dQw4w9WgXcQ',
...                    info_only=True)
site:                YouTube
title:               Rick Astley - Never Gonna Give You Up (Official Music Video)
stream:              # Best quality

    - itag:          22
      container:     MP4
      quality:       720p
      size:          15.2 MiB (15972320 bytes)

```

```python
>>> from you_get.extractors.pinterest import Pinterest
>>> pin = Pinterest()
>>> pin.download_by_url('https://www.pinterest.com/pin/123456789012345678/',
...                    info_only=True)
site:                Pinterest
title:               Cute cat picture
stream:              # Best quality

    - format:        original
      container:     jpg
      size:          0.8 MiB (839000 bytes)

```

Both examples utilize `download_by_url()` identically, even though the underlying implementations differ significantly.

## Key Files in the Extractor Architecture

| File | Role |
|------|------|
| **[`src/you_get/extractor.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractor.py)** | Defines `Extractor` and `VideoExtractor` base classes establishing the interface contract. |
| **[`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py)** | Provides utility functions (`get_content`, `match1`, `url_info`) and the dynamic extractor loader. |
| **`src/you_get/extractors/*.py`** | Site-specific implementations; each contains a class inheriting from `VideoExtractor` with `name`, `stream_types`, and overridden methods. |
| **`src/you_get/util/*.py`** | Low-level helpers for HTTP handling, logging, and JSON output used by extractors. |

## Summary

- **Site-specific extractors** are subclasses of `VideoExtractor` located in `src/you_get/extractors/`.
- **Mandatory implementation** requires overriding `prepare()` to fetch metadata and `extract()` to populate stream details.
- **Standard attributes** include `name` (site identifier) and `stream_types` (quality ordering).
- **Dynamic discovery** occurs automatically via [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) by inspecting the extractors package.
- **Complex sites** add internal helper methods while maintaining the same public interface as simple extractors.

## Frequently Asked Questions

### What methods must a site-specific extractor override?

A site-specific extractor must override **`prepare()`** to fetch the page and populate `self.title` and `self.streams`, and **`extract()`** to enrich each stream with container format, file size, and source URLs. All other methods, including `download_by_url()` and `download()`, inherit their behavior from `VideoExtractor`.

### How does you-get map URLs to extractors?

The framework uses dynamic discovery in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) to walk the `src/you_get/extractors/` directory, import each module, and match the URL host against the `sites` attribute of each extractor class. When matched, it instantiates the corresponding class automatically.

### What is the purpose of the stream_types attribute?

The `stream_types` attribute defines an ordered list of supported quality identifiers for a site. Each dictionary must contain either `id` or `itag` keys, and may specify `container`, `video_resolution`, and `quality`. The base class uses this list to sort available streams from highest to lowest quality.

### Can extractors implement additional helper methods?

Yes, extractors frequently implement private helper methods for site-specific logic such as signature decoding, API URL construction, or throttling removal. These helpers are internal implementation details and do not change the public interface, which remains consistent across all extractors through `download_by_url()`.