you-get Extractor Base Class Architecture and Site-Specific Inheritance Model
The you-get downloader implements a two-tier inheritance hierarchy in src/you_get/extractor.py where VideoExtractor extends the lightweight Extractor class to provide a standardized prepare → extract → download workflow, while concrete site extractors subclass VideoExtractor and override specific hooks to implement platform-specific media parsing.
The you-get project is a command-line utility for downloading media from various web sources. The extractor base class architecture centralizes common download logic while delegating site-specific parsing to specialized subclasses, enabling consistent handling of diverse video platforms through a unified interface.
Architecture of the you-get Extractor Base Class
The core extraction framework resides in src/you_get/extractor.py, which defines two fundamental classes that separate generic data storage from full download workflows.
The Two-Tier Class Hierarchy
The inheritance structure consists of distinct layers:
Extractor– A lightweight container class instantiated when users supply plain URLs. It holds generic attributes includingurl,title,vid,streams,dash_streams,caption_tracks,danmaku, andlyricsinitialized in__init__.VideoExtractor– ExtendsExtractorwith the complete download workflow implementation. This is the direct parent for all site-specific extractors in the codebase.
The VideoExtractor class defines the standardized pipeline that all concrete extractors follow, ensuring consistent behavior across different video platforms.
Core Workflow Methods and Attributes
The VideoExtractor class establishes several critical hooks and entry points:
Entry Points:
download_by_url(url, **kwargs)– Primary CLI entry point that accepts a target URLdownload_by_vid(vid, **kwargs)– Alternative entry point using video identifiers
Workflow Hooks:
prepare(self, **kwargs)– Optional setup hook for configuring headers, authentication, or proxy settings before extractionextract(self, **kwargs)– Mandatory override that performs HTTP requests, API parsing, and population ofself.streamsandself.dash_streamsdownload(self, **kwargs)– Common logic that selects optimal streams, constructs HTTP headers, and delegates todownload_urlsfromsrc/you_get/common.py
Utility Methods:
p(),p_stream(),p_i(),p_playlist()– Formatting helpers for human-readable output in info-only mode
How Site-Specific Extractors Inherit from VideoExtractor
Concrete implementations for individual websites follow a uniform subclassing pattern defined in files under src/you_get/extractors/.
Required Class Attributes and Overrides
Every site-specific extractor must define:
name– Class attribute providing the human-readable site identifier (e.g.,"YouTube","Bilibili") used by the printing utilitiesstream_types– List of dictionaries describing known format identifiers for non-DASH streams, enabling quality mapping in the download phaseextract(self, **kwargs)– Method implementation that fetches video metadata, parses API responses, and populates the instance attributes (self.title,self.streams,self.vid)
Optional overrides include:
prepare(self, **kwargs)– For site-specific authentication, cookie handling, or header configurationdownload_playlist_by_url()– For platforms supporting batch downloadscheck_playability_response()– For custom error handling when videos are region-blocked or removed
The Standard Implementation Pattern
Site extractors follow this consistent structure, as demonstrated in src/you_get/extractors/youtube.py:
from ..extractor import VideoExtractor
class YouTube(VideoExtractor):
name = "YouTube"
stream_types = [
# List of dicts mapping itags to quality descriptions
{'itag': '137', 'container': 'mp4', 'video_resolution': '1080p'},
# ... additional formats
]
def prepare(self, **kwargs):
# Configure headers, cookies, or API keys
self.headers = {'User-Agent': 'Mozilla/5.0 ...'}
def extract(self, **kwargs):
# 1. Fetch video page or API endpoint
# 2. Parse JSON or HTML for metadata
# 3. Populate stream dictionaries
self.title = "Parsed Video Title"
self.streams['hd'] = {
'url': 'https://...',
'container': 'mp4',
'size': 1024000,
'quality': 'HD'
}
Concrete Extractor Examples in the Codebase
The repository contains numerous implementations following this pattern:
src/you_get/extractors/youtube.py– Handles YouTube's complex itag system and signature decipheringsrc/you_get/extractors/bilibili.py– Implements Bilibili's API parsing and DASH stream handlingsrc/you_get/extractors/vimeo.py– Manages Vimeo's player configuration extractionsrc/you_get/extractors/acfun.py– Supports AcFun's specific video delivery formats
All files import VideoExtractor from ..extractor and override the extract() method to handle platform-specific JSON parsing, HTML scraping, or cryptographic signature handling.
Practical Usage of the Extractor Class
Developers can instantiate concrete extractors directly for programmatic access:
from you_get.extractors.youtube import YouTube
# Initialize and fetch metadata only
yt = YouTube()
yt.download_by_url('https://www.youtube.com/watch?v=dQw4w9WgXcQ',
info_only=True)
# Full download workflow
yt = YouTube()
yt.download_by_url('https://youtu.be/dQw4w9WgXcQ')
This invocation triggers the inherited workflow from VideoExtractor: prepare() configures the session, extract() populates the stream data, and download() handles the actual file retrieval using utilities from src/you_get/common.py and FFmpeg integration from src/you_get/processor/ffmpeg.py for DASH merging.
Summary
The you-get extractor base class architecture centralizes download orchestration while delegating platform-specific logic to subclasses:
VideoExtractorinsrc/you_get/extractor.pydefines theprepare→extract→downloadpipeline and common attributes likestreamsandtitle- Site-specific extractors inherit from
VideoExtractorand implement the mandatoryextract()method to parse site APIs - Required overrides include the
extract()method and class attributesnameandstream_types - Optional hooks like
prepare()enable custom authentication and header configuration - The design allows
download_by_url()anddownload_by_vid()to work uniformly across all supported platforms
Frequently Asked Questions
What is the difference between the Extractor and VideoExtractor classes in you-get?
The Extractor class is a lightweight container that holds generic attributes like url, title, and streams for basic URL handling. VideoExtractor extends this class in src/you_get/extractor.py to add the full download workflow including prepare(), extract(), and download() methods, making it the required base for all site-specific implementations.
Which methods must site-specific extractors override when inheriting from VideoExtractor?
Concrete extractors must override the extract(self, **kwargs) method to fetch and parse video metadata into self.streams and self.dash_streams. They must also define the class attributes name (site identifier string) and stream_types (format definitions). The prepare() method is optional but commonly overridden to set custom headers or authentication cookies.
How does you-get handle the download workflow across different video platforms?
The VideoExtractor class provides standardized entry points download_by_url() and download_by_vid() that execute a three-phase pipeline: prepare() for setup, extract() for metadata parsing, and download() for file retrieval. Site-specific extractors plug into this pipeline by overriding extract() to populate the stream dictionaries, while the base class handles stream selection, HTTP header construction, and delegation to download_urls() in src/you_get/common.py.
Where are concrete extractor implementations located in the you-get repository?
Site-specific extractors reside in src/you_get/extractors/, with each major platform having its own module (e.g., youtube.py, bilibili.py, vimeo.py). These modules import VideoExtractor from ..extractor and implement the required hooks to handle platform-specific APIs, HTML parsing, and authentication flows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →