you-get Extractor Base Class Architecture and Site-Specific Inheritance Model

The you-get downloader implements a two-tier inheritance hierarchy in src/you_get/extractor.py where VideoExtractor extends the lightweight Extractor class to provide a standardized prepare → extract → download workflow, while concrete site extractors subclass VideoExtractor and override specific hooks to implement platform-specific media parsing.

The you-get project is a command-line utility for downloading media from various web sources. The extractor base class architecture centralizes common download logic while delegating site-specific parsing to specialized subclasses, enabling consistent handling of diverse video platforms through a unified interface.

Architecture of the you-get Extractor Base Class

The core extraction framework resides in src/you_get/extractor.py, which defines two fundamental classes that separate generic data storage from full download workflows.

The Two-Tier Class Hierarchy

The inheritance structure consists of distinct layers:

  • Extractor – A lightweight container class instantiated when users supply plain URLs. It holds generic attributes including url, title, vid, streams, dash_streams, caption_tracks, danmaku, and lyrics initialized in __init__.
  • VideoExtractor – Extends Extractor with the complete download workflow implementation. This is the direct parent for all site-specific extractors in the codebase.

The VideoExtractor class defines the standardized pipeline that all concrete extractors follow, ensuring consistent behavior across different video platforms.

Core Workflow Methods and Attributes

The VideoExtractor class establishes several critical hooks and entry points:

Entry Points:

  • download_by_url(url, **kwargs) – Primary CLI entry point that accepts a target URL
  • download_by_vid(vid, **kwargs) – Alternative entry point using video identifiers

Workflow Hooks:

  • prepare(self, **kwargs) – Optional setup hook for configuring headers, authentication, or proxy settings before extraction
  • extract(self, **kwargs) – Mandatory override that performs HTTP requests, API parsing, and population of self.streams and self.dash_streams
  • download(self, **kwargs) – Common logic that selects optimal streams, constructs HTTP headers, and delegates to download_urls from src/you_get/common.py

Utility Methods:

  • p(), p_stream(), p_i(), p_playlist() – Formatting helpers for human-readable output in info-only mode

How Site-Specific Extractors Inherit from VideoExtractor

Concrete implementations for individual websites follow a uniform subclassing pattern defined in files under src/you_get/extractors/.

Required Class Attributes and Overrides

Every site-specific extractor must define:

  • name – Class attribute providing the human-readable site identifier (e.g., "YouTube", "Bilibili") used by the printing utilities
  • stream_types – List of dictionaries describing known format identifiers for non-DASH streams, enabling quality mapping in the download phase
  • extract(self, **kwargs) – Method implementation that fetches video metadata, parses API responses, and populates the instance attributes (self.title, self.streams, self.vid)

Optional overrides include:

  • prepare(self, **kwargs) – For site-specific authentication, cookie handling, or header configuration
  • download_playlist_by_url() – For platforms supporting batch downloads
  • check_playability_response() – For custom error handling when videos are region-blocked or removed

The Standard Implementation Pattern

Site extractors follow this consistent structure, as demonstrated in src/you_get/extractors/youtube.py:

from ..extractor import VideoExtractor

class YouTube(VideoExtractor):
    name = "YouTube"
    stream_types = [
        # List of dicts mapping itags to quality descriptions

        {'itag': '137', 'container': 'mp4', 'video_resolution': '1080p'},
        # ... additional formats

    ]

    def prepare(self, **kwargs):
        # Configure headers, cookies, or API keys

        self.headers = {'User-Agent': 'Mozilla/5.0 ...'}
        
    def extract(self, **kwargs):
        # 1. Fetch video page or API endpoint

        # 2. Parse JSON or HTML for metadata

        # 3. Populate stream dictionaries

        self.title = "Parsed Video Title"
        self.streams['hd'] = {
            'url': 'https://...',
            'container': 'mp4',
            'size': 1024000,
            'quality': 'HD'
        }

Concrete Extractor Examples in the Codebase

The repository contains numerous implementations following this pattern:

All files import VideoExtractor from ..extractor and override the extract() method to handle platform-specific JSON parsing, HTML scraping, or cryptographic signature handling.

Practical Usage of the Extractor Class

Developers can instantiate concrete extractors directly for programmatic access:

from you_get.extractors.youtube import YouTube

# Initialize and fetch metadata only

yt = YouTube()
yt.download_by_url('https://www.youtube.com/watch?v=dQw4w9WgXcQ', 
                   info_only=True)

# Full download workflow

yt = YouTube()
yt.download_by_url('https://youtu.be/dQw4w9WgXcQ')

This invocation triggers the inherited workflow from VideoExtractor: prepare() configures the session, extract() populates the stream data, and download() handles the actual file retrieval using utilities from src/you_get/common.py and FFmpeg integration from src/you_get/processor/ffmpeg.py for DASH merging.

Summary

The you-get extractor base class architecture centralizes download orchestration while delegating platform-specific logic to subclasses:

  • VideoExtractor in src/you_get/extractor.py defines the prepare → extract → download pipeline and common attributes like streams and title
  • Site-specific extractors inherit from VideoExtractor and implement the mandatory extract() method to parse site APIs
  • Required overrides include the extract() method and class attributes name and stream_types
  • Optional hooks like prepare() enable custom authentication and header configuration
  • The design allows download_by_url() and download_by_vid() to work uniformly across all supported platforms

Frequently Asked Questions

What is the difference between the Extractor and VideoExtractor classes in you-get?

The Extractor class is a lightweight container that holds generic attributes like url, title, and streams for basic URL handling. VideoExtractor extends this class in src/you_get/extractor.py to add the full download workflow including prepare(), extract(), and download() methods, making it the required base for all site-specific implementations.

Which methods must site-specific extractors override when inheriting from VideoExtractor?

Concrete extractors must override the extract(self, **kwargs) method to fetch and parse video metadata into self.streams and self.dash_streams. They must also define the class attributes name (site identifier string) and stream_types (format definitions). The prepare() method is optional but commonly overridden to set custom headers or authentication cookies.

How does you-get handle the download workflow across different video platforms?

The VideoExtractor class provides standardized entry points download_by_url() and download_by_vid() that execute a three-phase pipeline: prepare() for setup, extract() for metadata parsing, and download() for file retrieval. Site-specific extractors plug into this pipeline by overriding extract() to populate the stream dictionaries, while the base class handles stream selection, HTTP header construction, and delegation to download_urls() in src/you_get/common.py.

Where are concrete extractor implementations located in the you-get repository?

Site-specific extractors reside in src/you_get/extractors/, with each major platform having its own module (e.g., youtube.py, bilibili.py, vimeo.py). These modules import VideoExtractor from ..extractor and implement the required hooks to handle platform-specific APIs, HTML parsing, and authentication flows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →