# How the extractor_proxy Setting Routes Requests Through a Proxy in you-get

> Learn how you-get's extractor_proxy setting routes HTTP metadata requests through a proxy during extraction. Optimize your downloads by understanding this key feature.

- Repository: [Mort Yao/you-get](https://github.com/soimort/you-get)
- Tags: internals
- Published: 2026-03-06

---

**The `extractor_proxy` setting in you-get routes HTTP metadata requests through a specified proxy during the video extraction phase only, automatically removing the proxy before downloading actual media files.**

The `extractor_proxy` setting is a specialized networking feature in the [soimort/you-get](https://github.com/soimort/you-get) downloader that isolates proxy usage to the metadata discovery phase. This allows users to bypass regional restrictions or network filters when fetching video information while maintaining direct connections for the actual media download.

## What Is the extractor_proxy Setting?

`extractor_proxy` is exposed to users as the **`-y` / `--extractor-proxy`** command-line option in you-get. When specified, this setting instructs the tool to use an HTTP proxy **only while the extractor is running**—specifically during the phase where the program contacts the target website to discover video titles, available streams, and direct URLs.

Unlike a global proxy configuration that would route all traffic including large media files, this setting ensures that only the lightweight metadata requests pass through the proxy server. According to the source code in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) (lines 1652–1659), the CLI parser stores the proxy value in a global variable named `extractor_proxy`, making it available to the downloader functions.

## How extractor_proxy Routes Requests Through a Proxy

The routing mechanism involves three coordinated components across the codebase that temporarily install a proxy handler, execute the extraction logic, and then clean up the configuration before downloading begins.

### CLI Argument Parsing and Global Storage

In [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py), the argument parser defines the `-y` and `--extractor-proxy` flags during initialization. When provided, the proxy string (formatted as `HOST:PORT`) is parsed and stored globally. This value is later passed as a keyword argument to the extractor methods.

### Proxy Handler Implementation

The actual proxy installation logic resides in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) (lines 28–34) through two critical functions:

- **`set_proxy(host)`**: Constructs a `urllib.request.ProxyHandler` with the supplied host and port, then installs it as the default opener using `urllib.request.install_opener()`
- **`unset_proxy()`**: Removes the proxy configuration, restoring direct HTTP(S) connections

When `set_proxy` is called, **all subsequent HTTP(S) requests made via urllib are automatically routed through the specified proxy** until `unset_proxy` is invoked.

### Conditional Proxy Activation in Extractor Flow

The integration logic lives in [`src/you_get/extractor.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractor.py) within the `download_by_url` and `download_by_vid` methods. The implementation follows this precise sequence:

1. **Before extraction**: The code checks `if kwargs.get('extractor_proxy')` and executes `set_proxy(parse_host(kwargs['extractor_proxy']))` to enable the proxy
2. **During extraction**: The `self.prepare(**kwargs)` method runs, making all metadata requests through the proxy
3. **After extraction**: `unset_proxy()` is called immediately to remove the proxy before any media download starts

This guarantees that video stream URLs are fetched through the proxy (useful for geo-restricted sites), while the actual binary download occurs directly from the source server for optimal speed.

## Practical Usage Examples

Use the short flag to route extractor requests through a local proxy:

```bash
you-get -y 127.0.0.1:1080 https://www.youtube.com/watch?v=abc123

```

Or specify the long-form option with a remote proxy server:

```bash
you-get --extractor-proxy=proxy.example.com:8080 https://vimeo.com/987654

```

When using you-get as a Python library, manually control the proxy lifecycle:

```python
from you_get.common import set_proxy, unset_proxy, parse_host
from you_get.extractor import VideoExtractor

proxy = "127.0.0.1:1080"
ve = VideoExtractor()

# Enable proxy for extractor metadata phase only

set_proxy(parse_host(proxy))
ve.download_by_url('https://www.bilibili.com/video/av123456')

# Proxy automatically removed inside download_by_url before file download

```

## Summary

- **The `extractor_proxy` setting** (CLI flag `-y`) restricts proxy usage to the metadata extraction phase only
- **Source implementation** spans [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) (argument parsing and proxy handlers) and [`src/you_get/extractor.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractor.py) (conditional activation)
- **Proxy lifecycle** uses `set_proxy()` before `self.prepare()` and `unset_proxy()` immediately after, ensuring media downloads bypass the proxy
- **Technical mechanism** leverages `urllib.request.ProxyHandler` to intercept HTTP requests during the extraction window

## Frequently Asked Questions

### What is the difference between extractor_proxy and a global proxy in you-get?

The `extractor_proxy` setting routes only the initial metadata requests through a proxy, while a global proxy configuration would route all traffic including the actual video file downloads. According to the implementation in [`src/you_get/extractor.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractor.py), the proxy is explicitly unset before media download begins, keeping large transfers direct while only the extraction phase is proxied.

### When exactly is the proxy unset during the download process?

The proxy is unset immediately after the `self.prepare(**kwargs)` method completes in both `download_by_url` and `download_by_vid` methods within [`src/you_get/extractor.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractor.py). This occurs before the downloader fetches any video segments, ensuring the proxy is active solely during the metadata discovery phase.

### Can I use SOCKS5 proxies with extractor_proxy?

The current implementation in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) uses `urllib.request.ProxyHandler`, which primarily supports HTTP and HTTPS proxies. SOCKS5 support would require additional configuration of the `socks` library or `urllib.request` opener configuration, which is not implemented in the standard `set_proxy` function (lines 28–34).

### Where is the proxy configuration stored internally?

The proxy string is stored in the global variable `extractor_proxy` defined in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py), which is populated during argument parsing at lines 1652–1659. This value is passed through the `**kwargs` dictionary to extractor methods, where it is parsed by `parse_host()` and converted into a format suitable for `urllib.request.ProxyHandler`.