ReClip Architecture Explained: A Flask-Based YouTube Downloader Service

ReClip is a lightweight Flask web service that wraps yt-dlp to provide a REST API for video metadata extraction, playlist parsing, and media downloads through four architectural layers: request handling, business logic, process management, and storage delivery.

The ReClip architecture implemented in averygan/reclip offers a thin, container-friendly alternative to heavyweight download solutions. By combining Flask's HTTP capabilities with yt-dlp's extraction power, it delivers a stateless API suitable for ephemeral deployments without external database dependencies.

Four-Layer ReClip Architecture

The codebase organizes functionality into distinct layers, each with clear responsibilities and source file locations.

Request Handling Layer

Flask serves as the HTTP entry point, accepting JSON-encoded POST and GET requests.

  • Flask application initialization: app = Flask(__name__) in app.py lines 7–15
  • Route decorators: @app.route defines three main endpoints: /api/info, /api/playlist, and /api/download

This layer deserializes incoming JSON, validates required fields (url, optional format), and dispatches to business logic functions.

Business Logic Layer

Core operations reside in app.py lines 97–140, implementing three primary functions:

Function Purpose yt-dlp Command Equivalent
get_info Extract video metadata yt-dlp -j <url>
get_playlist_info List playlist entries yt-dlp --flat-playlist -J <url>
run_download Execute media download yt-dlp -f <format> -o <template> <url>

The parse_ytdlp_json helper processes yt-dlp's JSON output, selecting highest-bitrate formats per resolution for the formats array returned to clients.

Process Management Layer

ReClip handles long-running downloads through subprocess execution and background threading:


# From app.py lines 16-33 and 77-84

subprocess.run(
    cmd,
    capture_output=True,
    text=True,
    timeout=300
)

Key implementation details:

  • Job tracking: Global jobs dictionary stores in-memory state using UUID-based job_id keys
  • Threading: threading.Thread(daemon=True) spawns downloads without blocking the Flask worker
  • Status values: downloading, done, or error with descriptive messages

The timeout mechanism (300 seconds default) prevents hung processes from consuming resources indefinitely.

Storage and Delivery Layer

Downloaded media flows through a controlled lifecycle in app.py lines 34–75 and 99–105:

  1. Destination: DOWNLOAD_DIR constant (default: downloads/)
  2. Filename sanitization: Special characters removed to prevent path traversal
  3. Cleanup: Post-download removal of yt-dlp's intermediate .part files
  4. Delivery: Flask's send_file streams the final artifact with original filename preserved

No persistent storage beyond the filesystem—aligning with container-native deployment patterns.

ReClip Request Flow: Step by Step

Understanding how requests traverse the architecture clarifies design decisions.

Metadata Extraction Flow


POST /api/info → get_info() → subprocess.run(yt-dlp -j) 
→ parse_ytdlp_json() → JSON response with title, thumbnail, formats

The parse_ytdlp_json function specifically handles yt-dlp's verbose format objects, extracting height, vbr (video bitrate), abr (audio bitrate), and ext fields for client-side format selection.

Download Execution Flow


POST /api/download → start_download() → jobs[job_id] = {"status": "downloading"}
→ threading.Thread(target=run_download).start() → 202 Accepted with job_id

Clients then poll GET /api/status/<job_id> until status transitions to done, finally retrieving via GET /api/file/<job_id>.

Complete ReClip API Usage Example

import requests
import time

BASE_URL = "http://localhost:8899"

# 1. Inspect available formats

info_resp = requests.post(
    f"{BASE_URL}/api/info",
    json={"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ"}
)
info = info_resp.json()
print(f"Title: {info['title']}")
print(f"Available formats: {len(info['formats'])}")

# 2. Initiate download (video+audio merged)

download_resp = requests.post(
    f"{BASE_URL}/api/download",
    json={
        "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "format": "best"  # or specific format_id from info['formats']

    }
)
job_id = download_resp.json()["job_id"]

# 3. Poll for completion

while True:
    status_resp = requests.get(f"{BASE_URL}/api/status/{job_id}")
    status = status_resp.json()
    
    if status["status"] == "done":
        break
    elif status["status"] == "error":
        raise RuntimeError(f"Download failed: {status.get('error')}")
    
    time.sleep(1)

# 4. Retrieve final file

file_resp = requests.get(f"{BASE_URL}/api/file/{job_id}")
with open(status["filename"], "wb") as f:
    f.write(file_resp.content)

Key ReClip Source Files

File Lines of Code Architectural Role
app.py ~140 Complete server implementation: routing, logic, subprocess management
templates/index.html ~50 Browser-based UI demonstrating API consumption
requirements.txt 2 Flask dependency declaration
Dockerfile 12 Multi-stage build with yt-dlp binary installation
docker-compose.yml 8 Port 8899 exposure and volume mount

The entire service fits in a single Python file with no external service dependencies—a deliberate simplicity choice for the ReClip architecture.

Summary

  • ReClip implements a four-layer architecture: Flask request handling, yt-dlp business logic, threaded subprocess management, and filesystem storage delivery.

  • All state lives in the jobs dictionary—no database required, enabling horizontal scaling with load balancer sticky sessions or external job stores if needed.

  • Source code in app.py demonstrates production patterns: timeout handling, input sanitization, daemon threads, and streaming file responses.

  • The API design follows REST conventions with asynchronous job polling for long-running downloads, returning HTTP 202 on initiation and 200 on completion.

Frequently Asked Questions

What dependencies does ReClip require?

ReClip requires Flask (Python web framework) and yt-dlp (standalone binary). The requirements.txt lists only Flask; yt-dlp is installed via curl in the Dockerfile or must exist on the host PATH. No database, message queue, or cache systems are needed.

How does ReClip handle concurrent downloads?

Each download spawns an independent threading.Thread marked as daemon=True. Threads execute subprocess.run calls to yt-dlp, with status tracked in the global jobs dictionary. For production scale, deploy multiple ReClip containers behind a load balancer rather than increasing thread count.

Can ReClip download entire playlists?

Yes. The /api/playlist endpoint accepts playlist URLs and returns flat video entries using yt-dlp's --flat-playlist -J flags. Individual downloads must then be initiated separately via /api/download for each video URL in the playlist response.

Is ReClip suitable for production use?

The ReClip architecture targets containerized, short-lived deployments. In-memory job state means progress is lost on restart. For production, consider: externalizing jobs to Redis, adding authentication middleware, implementing rate limiting, and mounting DOWNLOAD_DIR to persistent or object storage rather than local filesystem.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →