How ReClip Manages Download Jobs: In-Memory Queue Architecture

ReClip handles video download jobs using an in-memory dictionary registry and daemon threads that execute yt-dlp commands asynchronously, preventing Flask request blocking while tracking job states from "downloading" to "done" or "error".

ReClip is a lightweight Flask-based video downloader that processes media requests without blocking the main web thread. Understanding how ReClip manages download jobs reveals a thread-based architecture that balances simplicity with performance for personal or small-scale deployments. The entire job management system is contained within the single-file Flask application in app.py.

Job Registry and State Storage

Global In-Memory Dictionary

At the core of ReClip's job management is a global jobs dictionary defined at line 13 in app.py:

jobs = {}

This registry stores every active and completed download job using randomly generated UUIDs as keys. Each job entry maintains a record containing the status, source url, video title, output file path, and sanitized filename for delivery.

Job Lifecycle States

Jobs transition through a simple finite state machine. When created, a job receives an initial status of "downloading". Upon completion of the background worker, the status updates to either "done" (success) or "error" (failure), with the latter populating an error key containing the exception or subprocess details.

Initiating Download Jobs

The Download Endpoint

Clients initiate downloads by POSTing to /api/download. According to the ReClip source code at lines 77-79, this endpoint:

  1. Generates a cryptographically random job_id
  2. Inserts a new entry into the jobs dictionary with the target URL and format preference
  3. Immediately returns the job_id to the client for polling

This non-blocking response ensures the HTTP connection closes within milliseconds, even for multi-gigabyte downloads.

Background Worker Threads

For each new job, ReClip spawns a dedicated daemon thread. Lines 80-82 in app.py implement this pattern:

thread = threading.Thread(target=run_download,
                          args=(job_id, url, format_choice, format_id))
thread.daemon = True
thread.start()

Marking the thread as daemon = True ensures that if the main Flask process terminates, all active downloads are automatically cleaned up without leaving zombie processes. The thread executes the run_download function, passing the job context and format selection arguments.

Download Execution and Error Handling

yt-dlp Command Construction

Inside run_download (lines 36-44), ReClip builds a yt-dlp command list based on the requested format:

  • Audio format: Downloads best audio and converts to MP3
  • Video format: Downloads best quality MP4 available

The function constructs the final command array and executes it via subprocess.run with a strict 5-minute timeout (line 48).

File Processing and Cleanup

After subprocess completion, lines 54-66 scan the downloads folder to locate the generated file. The logic differentiates between audio (*.mp3) and video (*.mp4) outputs, removes extraneous files generated during the download process, and sanitizes the filename for safe HTTP delivery.

State Transitions

Successful completions update the job record at lines 74-84:

job["status"] = "done"
job["file"] = chosen
job["filename"] = sanitized_name

Error handling (lines 50-53 and 85-89) captures non-zero exit codes, subprocess timeouts, and unexpected exceptions, setting job["status"] = "error" and populating job["error"] with diagnostic information for client visibility.

Status Retrieval and File Delivery

Polling Endpoint

Clients monitor progress via GET requests to /api/status/<job_id>. As implemented at lines 87-96, this endpoint returns a JSON object containing the current status, any error message, and the final filename once the job reaches the "done" state.

File Streaming

Once a job reports completion, the actual media file is retrieved via /api/file/<job_id>. Lines 99-104 use Flask's send_file to stream the stored file as an attachment, ensuring the browser prompts for download using the sanitized filename stored in the job registry.

Practical Implementation Examples

Starting a Download Job

import requests

resp = requests.post(
    "http://localhost:8899/api/download",
    json={"url": "https://www.youtube.com/watch?v=abc123",
          "format": "video"}  # or "audio"

)
job_id = resp.json()["job_id"]

Polling for Completion

import time
import requests

while True:
    status = requests.get(f"http://localhost:8899/api/status/{job_id}").json()
    print(status["status"])
    if status["status"] in ("done", "error"):
        break
    time.sleep(2)

Retrieving the Finished File

if status["status"] == "done":
    r = requests.get(f"http://localhost:8899/api/file/{job_id}", stream=True)
    with open(status["filename"], "wb") as f:
        for chunk in r.iter_content(8192):
            f.write(chunk)

Summary

  • In-Memory Registry: ReClip uses a global jobs dictionary in app.py (line 13) to track all download states without requiring external databases or message queues.
  • Thread-Based Concurrency: Each download spawns a daemon thread via threading.Thread (lines 80-82), keeping the Flask main thread responsive while yt-dlp processes media in the background.
  • Synchronous Subprocess Execution: The run_download function uses subprocess.run with a 300-second timeout to wrap yt-dlp, handling format conversion and file cleanup before updating job status.
  • Client Polling Architecture: Stateless HTTP endpoints (/api/status/<job_id> and /api/file/<job_id>) allow clients to poll for completion and stream results without persistent WebSocket connections.

Frequently Asked Questions

How does ReClip track the state of active download jobs?

ReClip tracks active downloads using the global jobs dictionary defined in app.py. Each entry maps a unique job ID to a dictionary containing status, url, title, file, and filename keys. The status field transitions from "downloading" to "done" or "error" as the background thread progresses, allowing the /api/status/<job_id> endpoint to report real-time state without database queries.

What happens if a download exceeds the 5-minute timeout?

If subprocess.run (line 48) exceeds the 300-second timeout, it raises a TimeoutExpired exception. The error handler at lines 85-89 catches this exception, updates the job status to "error", and stores the timeout details in the job["error"] field. The client polling the status endpoint will receive the error state and can handle the failure accordingly.

Is ReClip's job manager suitable for high-traffic production workloads?

No. ReClip's architecture uses an in-memory dictionary (jobs = {}) and spawns unlimited daemon threads per request, which limits scalability to the available RAM and thread count of the host machine. For production workloads requiring high concurrency, persistent storage, or crash recovery, the codebase would need migration to a proper task queue like Celery or RQ with Redis backing.

How are downloaded files cleaned up after retrieval?

ReClip performs cleanup during the run_download execution (lines 54-66) by removing extraneous files from the downloads folder, keeping only the target MP3 or MP4 file. However, the code does not automatically delete files after client download via /api/file/<job_id>—the file remains on disk until manually removed or overwritten by subsequent downloads with the same sanitized filename.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →