# How to Debug Issues with a Specific Site Extractor in you-get

> Debug you-get site extractor problems easily. Use the --debug flag and inspect extractor methods to find and fix regex misses, JSON errors, or decryption issues.

- Repository: [Mort Yao/you-get](https://github.com/soimort/you-get)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Enable the `--debug` flag to trace execution from the CLI through [`src/you_get/util/log.py`](https://github.com/soimort/you-get/blob/main/src/you_get/util/log.py) into the site-specific extractor located in `src/you_get/extractors/<site>.py`, then inspect the `prepare()` and `extract()` methods to identify regex misses, JSON parsing errors, or signature decryption failures.**

When a download fails for a specific site in the `you-get` repository, systematic debugging requires understanding how the extraction pipeline routes requests from the command-line interface to the site-specific implementation. The codebase provides built-in instrumentation through the [`util/log.py`](https://github.com/soimort/you-get/blob/main/util/log.py) module and supports interactive debugging via Python REPL, allowing you to isolate whether failures occur during initial page fetching, stream extraction, or signature deciphering.

## Understanding the you-get Extraction Architecture

The extraction framework follows a layered architecture that moves from generic CLI handling to site-specific logic:

- **CLI Entry Point**: [`src/you_get/__main__.py`](https://github.com/soimort/you-get/blob/main/src/you_get/__main__.py) parses arguments including `--debug` and delegates to [`main_dev.py`](https://github.com/soimort/you-get/blob/main/main_dev.py)
- **Logging Utility**: [`src/you_get/util/log.py`](https://github.com/soimort/you-get/blob/main/src/you_get/util/log.py) provides the `log.d()` debug wrapper that respects the `DEBUG` flag set via CLI
- **Network Layer**: [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) contains `get_content()` (line 463) and `post_content()` helpers that wrap HTTP requests
- **Base Extractors**: [`src/you_get/extractor.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractor.py) defines the `Extractor` and `VideoExtractor` base classes
- **Site Implementations**: `src/you_get/extractors/<site>.py` files (e.g., [`youtube.py`](https://github.com/soimort/you-get/blob/main/youtube.py), [`bilibili.py`](https://github.com/soimort/you-get/blob/main/bilibili.py)) implement `prepare()` and `extract()` methods

## Step-by-Step Debugging Workflow

### Enable Verbose Output with `--debug`

Start by running your command with the `--debug` flag to activate verbose logging throughout the stack:

```bash
you-get --debug "https://www.youtube.com/watch?v=dQw4w9WgXcQ"

```

This flag toggles `logging.debug` calls defined in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) around lines 1612-1616, enabling timestamped output from the [`util/log.py`](https://github.com/soimort/you-get/blob/main/util/log.py) logger.

### Trace the Execution Flow

The logger in [`src/you_get/util/log.py`](https://github.com/soimort/you-get/blob/main/src/you_get/util/log.py) provides a `d()` function that prints debug messages only when the debug level is enabled:

```python

# src/you_get/util/log.py

def d(msg, *args, **kwargs):
    if logging.getLogger().isEnabledFor(logging.DEBUG):
        print('[DEBUG]', msg % args, **kwargs)

```

When `--debug` is active, you will see output from the site extractor's `prepare()` method, such as `logging.debug('Extracting from the video page...')` in [`src/you_get/extractors/youtube.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractors/youtube.py) around lines 224-236.

### Locate the Site-Specific Extractor

The URL-to-extractor mapping resides in [`src/you_get/extractors/__init__.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractors/__init__.py). This registry imports classes like `YouTube` from [`youtube.py`](https://github.com/soimort/you-get/blob/main/youtube.py) and `BiliBili` from [`bilibili.py`](https://github.com/soimort/you-get/blob/main/bilibili.py), then returns the appropriate class via `match_extractor(url)`.

Identify your target site's module (e.g., [`src/you_get/extractors/youtube.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractors/youtube.py)) and examine its implementation of the `prepare()` and `extract()` methods.

### Inspect the `prepare()` Method

The `prepare()` method handles initial page fetching and initial data extraction. In [`src/you_get/extractors/youtube.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractors/youtube.py) around line 224, this method:

1. Fetches the video page using `get_content()` with custom headers
2. Extracts the `ytInitialPlayerResponse` JSON via regex
3. Retrieves the player JavaScript URL for signature deciphering

Look for debug statements surrounding these operations to verify HTTP success and regex matches. If the regex fails to find `ytInitialPlayerResponse`, the extraction stops before reaching stream parsing.

### Analyze the `extract()` Method

This method populates `self.streams` or `self.dash_streams` with downloadable URLs. Key debugging points include:

- **Signature Decryption**: Look for `s_to_sig` and `dethrottle` calls. YouTube frequently updates [`base.js`](https://github.com/soimort/you-get/blob/main/base.js), breaking these functions. Check the comment block around lines 78-119 in [`youtube.py`](https://github.com/soimort/you-get/blob/main/youtube.py) for the current deciphering logic.
- **Stream Validation**: Verify each stream entry contains a usable URL. Missing signature data triggers warnings like `log.wtf('No signatureCipher ...')`.

### Verify Network Requests

If HTTP requests fail, examine [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) where `get_content()` and `get_https()` reside. These functions accept a `debuglevel` parameter that passes to `conn.set_debuglevel()`:

```python

# Temporarily modify the call in prepare()

self.js = get_content(self.html5player, debuglevel=1).replace('\n', ' ')

```

Setting `debuglevel=1` prints raw request/response headers to stderr, revealing TLS errors, redirects, or 403 Forbidden responses.

### Add Custom Debug Statements

When built-in logs are insufficient, insert temporary debug lines using the utility logger:

```python
from you_get.util import log
log.d('Current VID: %s', self.vid)
log.d('Player URL: %s', self.html5player)

```

Re-run with `--debug` to see these values without modifying global logging configuration.

### Test the Extractor in Isolation

Bypass the CLI wrapper to test logic directly in a Python REPL:

```python
>>> from you_get.extractors.youtube import YouTube
>>> yt = YouTube('https://youtu.be/dQw4w9WgXcQ')
>>> yt.prepare()
>>> yt.extract()
>>> yt.streams  # Inspect the generated URL dictionary

```

This approach eliminates argument parsing overhead and allows immediate inspection of `yt.streams`, `yt.title`, and intermediate variables.

## Common Failure Points and Solutions

When debugging specific extractors, three issues occur most frequently:

- **Regex Pattern Misses**: Sites update their HTML structure, breaking patterns used in `match1()` (defined in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py)). Update the regex to match new page layouts.
- **Player Code Changes**: YouTube's [`base.js`](https://github.com/soimort/you-get/blob/main/base.js) updates invalidate the `s_to_sig` deciphering algorithm. Monitor the player script URL logged in debug mode and update the transformation logic accordingly.
- **Missing JSON Fields**: API responses may exclude expected keys. Add defensive `try/except` blocks around `json.loads()` calls and log the raw response for analysis.

## Summary

- **Enable `--debug`** to activate the [`util/log.py`](https://github.com/soimort/you-get/blob/main/util/log.py) logger and see execution flow through [`src/you_get/__main__.py`](https://github.com/soimort/you-get/blob/main/src/you_get/__main__.py).
- **Locate the extractor** in `src/you_get/extractors/<site>.py` using the registry in [`__init__.py`](https://github.com/soimort/you-get/blob/main/__init__.py).
- **Inspect `prepare()`** for HTTP failures in `get_content()` calls and regex extraction of JSON/player data.
- **Check `extract()`** for signature decryption errors in `s_to_sig` and proper population of `self.streams`.
- **Use `debuglevel=1`** in `get_content()` calls to inspect raw HTTP headers when network issues occur.
- **Test interactively** via Python REPL by importing the extractor class directly to bypass CLI overhead.

## Frequently Asked Questions

### How do I enable debug mode in you-get?

Run any command with the `--debug` flag. This sets the logging level to `DEBUG` in [`src/you_get/common.py`](https://github.com/soimort/you-get/blob/main/src/you_get/common.py) (lines 1612-1616) and activates verbose output from [`src/you_get/util/log.py`](https://github.com/soimort/you-get/blob/main/src/you_get/util/log.py), showing HTTP requests, regex matches, and stream extraction steps.

### Where are the site-specific extractors located?

Site extractors reside in `src/you_get/extractors/` as individual Python modules (e.g., [`youtube.py`](https://github.com/soimort/you-get/blob/main/youtube.py), [`bilibili.py`](https://github.com/soimort/you-get/blob/main/bilibili.py)). The registry in [`src/you_get/extractors/__init__.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractors/__init__.py) maps URL patterns to these classes. Each module implements `prepare()` to fetch metadata and `extract()` to populate stream URLs.

### What should I check if a YouTube video fails to download?

First, verify the `prepare()` method in [`src/you_get/extractors/youtube.py`](https://github.com/soimort/you-get/blob/main/src/you_get/extractors/youtube.py) successfully retrieves the page and extracts `ytInitialPlayerResponse`. If that passes, check the `s_to_sig` function (lines 78-119) for signature decryption errors, as YouTube updates [`base.js`](https://github.com/soimort/you-get/blob/main/base.js) frequently. Enable `--debug` to see which specific step fails.

### How can I test an extractor without using the CLI?

Import the extractor class directly into a Python REPL or script. For example, `from you_get.extractors.youtube import YouTube`, instantiate with a URL, and call `prepare()` followed by `extract()`. This allows inspection of `self.streams` and intermediate variables without CLI argument parsing overhead.