# ReClip Architecture Explained: A Flask-Based YouTube Downloader Service

> Discover the ReClip architecture, a Flask-based YouTube downloader service. Explore its four layers: request handling, business logic, process management, and storage delivery for efficient media downloads.

- Repository: [Avery Gan/reclip](https://github.com/averygan/reclip)
- Tags: architecture
- Published: 2026-09-03

---

**ReClip is a lightweight Flask web service that wraps yt-dlp to provide a REST API for video metadata extraction, playlist parsing, and media downloads through four architectural layers: request handling, business logic, process management, and storage delivery.**

The **ReClip architecture** implemented in [averygan/reclip](https://github.com/averygan/reclip) offers a thin, container-friendly alternative to heavyweight download solutions. By combining Flask's HTTP capabilities with yt-dlp's extraction power, it delivers a stateless API suitable for ephemeral deployments without external database dependencies.

## Four-Layer ReClip Architecture

The codebase organizes functionality into distinct layers, each with clear responsibilities and source file locations.

### Request Handling Layer

Flask serves as the HTTP entry point, accepting JSON-encoded POST and GET requests.

- **Flask application initialization**: `app = Flask(__name__)` in [`app.py`](https://github.com/averygan/reclip/blob/main/app.py) lines 7–15
- **Route decorators**: `@app.route` defines three main endpoints: `/api/info`, `/api/playlist`, and `/api/download`

This layer deserializes incoming JSON, validates required fields (`url`, optional `format`), and dispatches to business logic functions.

### Business Logic Layer

Core operations reside in [`app.py`](https://github.com/averygan/reclip/blob/main/app.py) lines 97–140, implementing three primary functions:

| Function | Purpose | yt-dlp Command Equivalent |
|----------|---------|---------------------------|
| `get_info` | Extract video metadata | `yt-dlp -j <url>` |
| `get_playlist_info` | List playlist entries | `yt-dlp --flat-playlist -J <url>` |
| `run_download` | Execute media download | `yt-dlp -f <format> -o <template> <url>` |

The `parse_ytdlp_json` helper processes yt-dlp's JSON output, selecting highest-bitrate formats per resolution for the `formats` array returned to clients.

### Process Management Layer

ReClip handles long-running downloads through **subprocess execution** and **background threading**:

```python

# From app.py lines 16-33 and 77-84

subprocess.run(
    cmd,
    capture_output=True,
    text=True,
    timeout=300
)

```

Key implementation details:

- **Job tracking**: Global `jobs` dictionary stores in-memory state using UUID-based `job_id` keys
- **Threading**: `threading.Thread(daemon=True)` spawns downloads without blocking the Flask worker
- **Status values**: `downloading`, `done`, or `error` with descriptive messages

The timeout mechanism (300 seconds default) prevents hung processes from consuming resources indefinitely.

### Storage and Delivery Layer

Downloaded media flows through a controlled lifecycle in [`app.py`](https://github.com/averygan/reclip/blob/main/app.py) lines 34–75 and 99–105:

1. **Destination**: `DOWNLOAD_DIR` constant (default: `downloads/`)
2. **Filename sanitization**: Special characters removed to prevent path traversal
3. **Cleanup**: Post-download removal of yt-dlp's intermediate `.part` files
4. **Delivery**: Flask's `send_file` streams the final artifact with original filename preserved

No persistent storage beyond the filesystem—aligning with container-native deployment patterns.

## ReClip Request Flow: Step by Step

Understanding how requests traverse the architecture clarifies design decisions.

### Metadata Extraction Flow

```

POST /api/info → get_info() → subprocess.run(yt-dlp -j) 
→ parse_ytdlp_json() → JSON response with title, thumbnail, formats

```

The `parse_ytdlp_json` function specifically handles yt-dlp's verbose format objects, extracting `height`, `vbr` (video bitrate), `abr` (audio bitrate), and `ext` fields for client-side format selection.

### Download Execution Flow

```

POST /api/download → start_download() → jobs[job_id] = {"status": "downloading"}
→ threading.Thread(target=run_download).start() → 202 Accepted with job_id

```

Clients then poll `GET /api/status/<job_id>` until status transitions to `done`, finally retrieving via `GET /api/file/<job_id>`.

## Complete ReClip API Usage Example

```python
import requests
import time

BASE_URL = "http://localhost:8899"

# 1. Inspect available formats

info_resp = requests.post(
    f"{BASE_URL}/api/info",
    json={"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ"}
)
info = info_resp.json()
print(f"Title: {info['title']}")
print(f"Available formats: {len(info['formats'])}")

# 2. Initiate download (video+audio merged)

download_resp = requests.post(
    f"{BASE_URL}/api/download",
    json={
        "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "format": "best"  # or specific format_id from info['formats']

    }
)
job_id = download_resp.json()["job_id"]

# 3. Poll for completion

while True:
    status_resp = requests.get(f"{BASE_URL}/api/status/{job_id}")
    status = status_resp.json()
    
    if status["status"] == "done":
        break
    elif status["status"] == "error":
        raise RuntimeError(f"Download failed: {status.get('error')}")
    
    time.sleep(1)

# 4. Retrieve final file

file_resp = requests.get(f"{BASE_URL}/api/file/{job_id}")
with open(status["filename"], "wb") as f:
    f.write(file_resp.content)

```

## Key ReClip Source Files

| File | Lines of Code | Architectural Role |
|------|---------------|---------------------|
| [`app.py`](https://github.com/averygan/reclip/blob/main/app.py) | ~140 | Complete server implementation: routing, logic, subprocess management |
| [`templates/index.html`](https://github.com/averygan/reclip/blob/main/templates/index.html) | ~50 | Browser-based UI demonstrating API consumption |
| [`requirements.txt`](https://github.com/averygan/reclip/blob/main/requirements.txt) | 2 | Flask dependency declaration |
| `Dockerfile` | 12 | Multi-stage build with yt-dlp binary installation |
| [`docker-compose.yml`](https://github.com/averygan/reclip/blob/main/docker-compose.yml) | 8 | Port 8899 exposure and volume mount |

The entire service fits in a single Python file with no external service dependencies—a deliberate simplicity choice for the **ReClip architecture**.

## Summary

- **ReClip** implements a four-layer architecture: Flask request handling, yt-dlp business logic, threaded subprocess management, and filesystem storage delivery.

- All state lives in the **`jobs`** dictionary—no database required, enabling horizontal scaling with load balancer sticky sessions or external job stores if needed.

- Source code in [`app.py`](https://github.com/averygan/reclip/blob/main/app.py) demonstrates production patterns: timeout handling, input sanitization, daemon threads, and streaming file responses.

- The API design follows REST conventions with asynchronous job polling for long-running downloads, returning HTTP 202 on initiation and 200 on completion.

## Frequently Asked Questions

### What dependencies does ReClip require?

ReClip requires **Flask** (Python web framework) and **yt-dlp** (standalone binary). The [`requirements.txt`](https://github.com/averygan/reclip/blob/main/requirements.txt) lists only Flask; yt-dlp is installed via `curl` in the Dockerfile or must exist on the host PATH. No database, message queue, or cache systems are needed.

### How does ReClip handle concurrent downloads?

Each download spawns an independent `threading.Thread` marked as `daemon=True`. Threads execute `subprocess.run` calls to yt-dlp, with status tracked in the global `jobs` dictionary. For production scale, deploy multiple ReClip containers behind a load balancer rather than increasing thread count.

### Can ReClip download entire playlists?

Yes. The `/api/playlist` endpoint accepts playlist URLs and returns flat video entries using yt-dlp's `--flat-playlist -J` flags. Individual downloads must then be initiated separately via `/api/download` for each video URL in the playlist response.

### Is ReClip suitable for production use?

The **ReClip architecture** targets containerized, short-lived deployments. In-memory job state means progress is lost on restart. For production, consider: externalizing `jobs` to Redis, adding authentication middleware, implementing rate limiting, and mounting `DOWNLOAD_DIR` to persistent or object storage rather than local filesystem.