How you-get Implements Network Request Handling and Error Recovery

The you-get tool centralizes all HTTP communication in src/you_get/common.py, providing a unified retry mechanism that attempts requests up to three times while handling timeouts, HTTP errors, SSL flexibility, and content compression automatically.

The open-source video downloader soimort/you-get streamlines network resilience through a single module that every site-specific extractor depends on. By funneling all requests through a centralized retry wrapper, the codebase ensures consistent error recovery whether downloading from Bilibili, Vimeo, or any other supported platform.

Core Network Architecture in common.py

All network operations in you-get flow through src/you_get/common.py, which exposes a small, well-defined API used by extractor scripts across the codebase.

The Retry Wrapper: urlopen_with_retry

The foundation of you-get's resilience is urlopen_with_retry(*args, **kwargs), a thin wrapper around urllib.request.urlopen. According to the source code, this function implements a hardcoded retry strategy:

  • It attempts the request up to three times
  • It catches socket.timeout and urllib.error.HTTPError
  • On the final failure, it re-raises the exception to signal permanent failure

This wrapper ensures that transient network timeouts and temporary CDN failures do not immediately crash the download process.

High-Level HTTP Helpers

Building on the retry wrapper, common.py provides convenience functions that handle headers, cookies, compression, and encoding:

  • get_content(url, headers={}, decoded=True) — Sends a GET request and returns the response body as a string. It automatically injects a fake User-Agent header, manages cookies, detects charset from the Content-Type header, and decompresses gzip/deflate payloads.

  • post_content(url, headers={}, post_data={}, decoded=True, **kwargs) — Handles POST requests with the same retry logic and header management as get_content. Supports both form-encoded data and raw payloads via post_data_raw.

  • get_location(url, headers=None, get_method='HEAD') — Performs HEAD requests (or other methods) to resolve final redirected URLs, inheriting the retry behavior of urlopen_with_retry.

Error Recovery and Resilience Mechanisms

The you-get network stack implements multiple layers of error recovery beyond simple retries.

Automatic Retry Logic for Transient Failures

Every network call funnels through urlopen_with_retry, guaranteeing consistent handling of intermittent failures. When a socket.timeout occurs, the function pauses and retries. For urllib.error.HTTPError, it logs the HTTP status code before attempting recovery, providing visibility into CDN or server-side issues.

SSL Flexibility for Self-Signed Certificates

When the global insecure flag is set, you-get creates an SSL context with verify_mode = ssl.CERT_NONE in urlopen_with_retry. This allows connections to servers with self-signed or expired certificates without modifying system certificate stores.

If a http.cookiejar.CookieJar instance is populated, the utilities manually inject a Cookie header into requests. This bypasses limitations in Python versions below 3.10 that could not correctly handle HttpOnly cookies through the standard cookie processor.

Compression and Encoding Recovery

After receiving responses, the utilities automatically decompress gzip or deflate payloads using helper functions (ungzip, undeflate). For character encoding, get_content and post_content inspect the Content-Type header for charset declarations; if none exists, they fall back to UTF-8 to prevent decoding errors.

Usage in Extractor Modules

All site-specific extractors import these centralized helpers rather than implementing their own network logic. For example, you_get/extractors/bilibili.py and you_get/extractors/vimeo.py import:

from ..common import get_content, post_content, urlopen_with_retry

This architectural choice means that improvements to retry logic, header spoofing, or SSL handling in common.py propagate immediately to every supported video platform without modifying individual extractor code.

Code Examples

Simple GET with Automatic Retries

Retrieve video page HTML with automatic charset detection and decompression:

from you_get.common import get_content

html = get_content('https://example.com/video/12345')
print(html[:200])  # First 200 characters

POST Request with Form Data

Submit login credentials with proper headers and automatic retry handling:

from you_get.common import post_content

payload = {'username': 'myuser', 'password': 'secret'}
response = post_content(
    'https://example.com/api/login',
    headers={'User-Agent': 'you-get/2.4.0'},
    post_data=payload,
    decoded=True
)
print(response)

Resolving Redirected URLs

Determine the final destination of a shortened URL using a HEAD request:

from you_get.common import get_location

final_url = get_location('http://short.url/abc')
print('Redirected to:', final_url)

Manual Retry Wrapper Usage

Access the underlying retry mechanism directly for custom request configurations:

from urllib.request import Request
from you_get.common import urlopen_with_retry

req = Request('https://example.com/api')
resp = urlopen_with_retry(req, timeout=10)  # Up to 3 retries on timeout/HTTPError

data = resp.read()

Summary

  • Centralized network layer: All HTTP operations in you-get flow through src/you_get/common.py, ensuring consistent behavior across extractors.
  • Triple-retry strategy: urlopen_with_retry catches socket.timeout and HTTPError, attempting requests up to three times before failing.
  • Compression and encoding: Automatic handling of gzip/deflate decompression and charset detection from Content-Type headers.
  • SSL bypass capability: Support for insecure connections via ssl.CERT_NONE when the insecure flag is enabled.
  • Cookie compatibility: Manual header injection to support HttpOnly cookies in older Python versions.

Frequently Asked Questions

How many times does you-get retry a failed network request?

The urlopen_with_retry function in src/you_get/common.py attempts each request up to three times. It catches socket.timeout and urllib.error.HTTPError on the first two attempts, re-raising the exception only on the final failure to prevent infinite loops.

Can you-get handle websites with self-signed SSL certificates?

Yes. When the global insecure flag is set, you-get creates an SSL context with verify_mode = ssl.CERT_NONE inside urlopen_with_retry. This disables certificate verification, allowing connections to servers with self-signed or invalid certificates.

Where is the User-Agent header defined in you-get?

The fake User-Agent headers used by get_content and post_content are defined in src/you_get/util/term.py as fake_headers. These headers are automatically injected into requests to mimic browser behavior and prevent blocking by video platforms.

Do all video extractors use the same network error handling?

Yes. All site-specific extractors in src/you_get/extractors/ (such as bilibili.py and vimeo.py) import get_content, post_content, and urlopen_with_retry from common.py. This ensures that any improvements to retry logic, timeout handling, or compression support automatically benefit every supported video platform.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →