How the JSON Output Mode Serializes Download Information in you-get
When you run you-get with --json, the tool constructs a fake VideoExtractor object containing metadata and stream URLs, then serializes specific fields—url, title, site, streams, and optional extra—into a JSON object via json_output.output() instead of downloading files.
The soimort/you-get command-line utility provides a JSON output mode that extracts video metadata and download URLs without performing actual downloads. This feature routes the download workflow through a dedicated serialization module that converts internal VideoExtractor objects into structured JSON data suitable for scripting and automation.
How the JSON Flag Redirects the Download Workflow
In src/you_get/common.py, the command-line argument --json sets an internal json_output flag at lines 993–996. When this flag is True, you-get bypasses the standard printing and downloading logic entirely and delegates output handling to the src/you_get/json_output.py module.
- If
json_outputis enabled, calls to download functions are intercepted. - The
json_output.download_urls()function creates a temporary extractor object rather than initiating file transfers. - All JSON-specific logic remains isolated in
json_output.py, leaving the core extractor code unchanged.
Constructing the VideoExtractor Object for JSON Output
The json_output.download_urls() function (lines 49–66 in src/you_get/json_output.py) builds a fake VideoExtractor instance to hold download metadata. This object mimics the structure used during actual downloads but contains only serialization-relevant fields:
ve = last_info # Reused from previous print_info call or newly created
stream = {}
stream['container'] = ext # File extension (e.g., 'mp4')
stream['size'] = total_size # Size in bytes (may be None)
stream['src'] = urls # List of real URLs to download
if refer:
stream['refer'] = refer
stream['video_profile'] = '__default__'
ve.streams = {'__default__': stream}
output(ve) # Serialize to JSON
This temporary structure captures the container type, file size, source URLs, and referer information that the regular download workflow would use, packaging it into the ve.streams dictionary under the __default__ key.
Serializing Metadata to JSON
The json_output.output() function (lines 7–34 in src/you_get/json_output.py) converts the VideoExtractor into a plain Python dictionary and emits it via json.dumps(). The serialization captures both mandatory and optional fields:
Base fields always included:
url– The original video URL(s) fromve.urltitle– The extracted video title fromve.titlesite– The site identifier fromve.namestreams– The complete mapping of quality profiles to stream data
Optional fields added when present:
audiolang– Included ifve.audiolangexists on the extractorextra– A nested object containingrefererand/orua(user-agent) strings when available
The output format depends on the pretty_print argument: if True, the JSON uses indent=4 for readability; otherwise, it outputs compact, single-line JSON.
JSON Schema and Structure
A typical JSON output from you-get follows this deterministic structure:
{
"url": "https://example.com/video/12345",
"title": "Sample Video",
"site": "youtube",
"streams": {
"__default__": {
"container": "mp4",
"size": 12345678,
"src": ["https://cdn.example.com/vid12345.mp4"],
"video_profile": "__default__"
}
},
"extra": {
"referer": "https://example.com/",
"ua": "Mozilla/5.0 …"
}
}
The streams object may contain multiple keys beyond __default__ if the specific site extractor populates different quality profiles (e.g., 720p, 1080p) before calling the JSON output functions. Each stream entry maintains the same schema: container, size, src, and video_profile.
Info-Only Mode vs Full Download Metadata
When --json is combined with --info (or -i), you-get invokes json_output.print_info() at lines 40–47 instead of download_urls(). This creates a minimal VideoExtractor containing only the site and title fields, with no streams data.
The resulting JSON contains only:
urltitlesite
This mode is useful for quickly identifying videos and checking availability without fetching stream URLs or file metadata.
Practical Usage Examples
Generate complete download metadata including stream URLs:
you-get --json "https://www.youtube.com/watch?v=abcd1234"
Fetch only basic video information without stream details:
you-get -i --json "https://vimeo.com/567890"
Both commands output valid JSON to stdout, with the first example including the full streams array containing actual CDN URLs ready for external download tools.
Summary
- Flag handling: The
--jsonargument insrc/you_get/common.py(lines 993–996) redirects output to the JSON serialization module. - Object construction:
json_output.download_urls()builds a temporaryVideoExtractorwith__default__stream data includingcontainer,size, andsrcURLs. - Serialization:
json_output.output()converts the extractor to a dictionary with fieldsurl,title,site,streams, and optionalaudiolang/extrametadata. - Dual modes: Full metadata mode includes stream URLs, while info-only mode (
-i --json) outputs only title and site identification. - Implementation: All JSON logic is isolated in
src/you_get/json_output.py, maintaining clean separation from download and extraction code.
Frequently Asked Questions
What fields are included in you-get's JSON output?
The JSON object always includes url, title, site, and streams. The streams object maps quality profiles to objects containing container (file extension), size (bytes), src (URL list), and video_profile. Optional fields include audiolang for multi-language videos and extra (containing referer and ua) when the extractor provides HTTP headers.
How do I enable JSON output mode in you-get?
Append the --json flag to any you-get command. For metadata-only output without stream URLs, combine it with --info or -i. According to the source code in src/you_get/common.py, this sets the internal json_output flag and routes calls through src/you_get/json_output.py instead of the standard download pipeline.
Why does the JSON output use __default__ as a stream key?
The json_output.download_urls() function creates a single-entry streams dictionary with the key __default__ to maintain compatibility with the VideoExtractor structure used throughout you-get. If a site extractor populates multiple quality profiles before JSON serialization, additional keys (e.g., 720p, 1080p) will appear alongside __default__.
Where is the JSON serialization logic located?
All JSON-specific functionality resides in src/you_get/json_output.py. This module contains three primary functions: output() (handles the actual json.dumps conversion), download_urls() (builds the fake extractor for full metadata), and print_info() (creates minimal extractors for info-only mode). The module is called conditionally from src/you_get/common.py when the --json flag is detected.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →