How the FastAPI WebUI Integrates with the MediaCrawler Core Engine

The FastAPI WebUI in MediaCrawler acts as a thin HTTP control plane that manages the crawling engine through a dedicated service layer, handling process lifecycle, log streaming, and status monitoring via REST endpoints and WebSockets.

The MediaCrawler open-source project provides a browser-based interface for controlling multi-platform content crawlers. Rather than embedding the crawling logic directly into the web server, the FastAPI WebUI operates as a separate orchestration layer that spawns and monitors the core crawler as an asynchronous subprocess, enabling safe concurrent operation and real-time log streaming.

FastAPI Application Structure

Main Application Setup

The FastAPI application is initialized in api/main.py, where the server registers CORS middleware and mounts the compiled Vue/Vite frontend bundle. Static assets are served under /assets, /logos, and /static, while the root endpoint (GET /) returns index.html to bootstrap the single-page application.


# From api/main.py - FastAPI instantiation and static mounting

app = FastAPI()
app.mount("/assets", StaticFiles(directory="api/webui/assets"), name="assets")

Router Architecture

The API routes are organized into three specialized routers defined in api/routers/__init__.py:

  • crawler_router – Controls crawler lifecycle (/api/crawler/start, /api/crawler/stop, /api/crawler/status, /api/crawler/logs)
  • data_router – Handles data export operations (JSON, CSV, Excel downloads)
  • websocket_router – Provides real-time log streaming capabilities

These routers are included with a common /api prefix, creating a clean separation between the web interface and the underlying crawling engine.

Crawler Control Endpoints

The crawler_router (implemented in api/routers/crawler.py) delegates all operations to the crawler manager service rather than executing crawl logic directly. This design prevents the FastAPI event loop from blocking during long-running scraping operations.

Starting the Crawler

When the frontend sends a POST request to /api/crawler/start, the endpoint receives a CrawlerStartRequest Pydantic model containing platform selection, crawling mode, keywords, and authentication options. The router passes this to crawler_manager.start(request), which launches the actual crawler entry point (main.py) as an async subprocess using asyncio.create_subprocess_exec.


# Conceptual flow from api/routers/crawler.py

@crawler_router.post("/start")
async def start_crawler(request: CrawlerStartRequest):
    return await crawler_manager.start(request)

Stopping and Monitoring

The control endpoints provide safe process management:

  • POST /api/crawler/stop – Invokes crawler_manager.stop(), which terminates the subprocess cleanly using process.terminate() and monitors the exit code
  • GET /api/crawler/status – Returns a CrawlerStatusResponse containing the running state, process ID (pid), start timestamp, and current log line count by checking process.poll()
  • GET /api/crawler/logs?limit=n – Retrieves the most recent entries from the manager's circular log buffer
// Polling crawler status from the frontend
setInterval(async () => {
  const resp = await fetch('/api/crawler/status');
  const status = await resp.json();
  console.log('Crawler running:', status.running, 'PID:', status.pid);
}, 2000);

The Crawler Manager Service

The crawler_manager.py service encapsulates all interaction with the crawling engine, maintaining singleton process state and providing concurrency safety.

Process Management

The service maintains an internal process reference that tracks the running crawler subprocess. Before starting a new crawl, it checks process.poll() to ensure no existing process is running, returning HTTP 400 if a duplicate start request is detected. The manager constructs the command-line invocation using uv run main.py with the parameters extracted from the CrawlerStartRequest model.

Log Collection and Status Tracking

While the crawler runs, the manager asynchronously reads from subprocess stdout and stderr, parsing each line into a CrawlerLog Pydantic model with timestamp, log level, and message fields. These entries are stored in a bounded list for quick retrieval by the logs endpoint. The CrawlerStatusResponse schema exposed by the service provides the frontend with accurate runtime metadata including running, pid, started_at, and log_line_count.


# From api/services/crawler_manager.py

class CrawlerStatusResponse(BaseModel):
    running: bool
    pid: int | None
    started_at: datetime | None
    log_line_count: int

Frontend Integration

The Vue-based frontend (bundled in api/webui) communicates with the FastAPI layer through standard HTTP requests and WebSocket connections. When initiating a crawl, the UI serializes the configuration into a CrawlerStartRequest JSON payload:

await fetch('/api/crawler/start', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    platform: 'dy',
    mode: 'search',
    keyword: 'funny cats',
    save_option: 'json',
    login_type: 'qrcode'
  })
});

For real-time updates, the frontend either polls /api/crawler/status and /api/crawler/logs periodically or connects to the WebSocket endpoint for push-based log streaming. This architecture decouples the heavy I/O of web scraping from the web server's event loop while keeping both components within the same Python process space for easy deployment.

Summary

  • FastAPI WebUI serves as a control plane in api/main.py, mounting static frontend assets and registering route prefixes
  • Three dedicated routers (crawler_router, data_router, websocket_router) handle distinct concerns under the /api prefix
  • Crawler manager service (api/services/crawler_manager.py) encapsulates subprocess lifecycle, using asyncio.create_subprocess_exec to launch main.py and monitoring via process.poll()
  • Pydantic models (CrawlerStartRequest, CrawlerStatusResponse, CrawlerLog) enforce type safety between the UI and crawling engine
  • Real-time log streaming is achieved through asynchronous stdout/stderr reading into a bounded buffer, accessible via REST or WebSocket endpoints

Frequently Asked Questions

How does the FastAPI WebUI prevent starting multiple crawler instances simultaneously?

The crawler manager service maintains a singleton process reference and checks process.poll() before accepting new start requests. If a process is already running, it returns HTTP 400 with a clear error message indicating that a crawl is currently active, ensuring safe concurrency control at the application layer.

What happens to crawler logs when the WebUI is not actively polling?

The crawler manager stores log entries in an in-memory circular buffer (bounded list) while the subprocess runs. Even if the frontend disconnects or stops polling, logs continue accumulating in this buffer up to the configured limit. When the client resumes polling /api/crawler/logs, it retrieves the most recent entries from this buffer, including those generated during the disconnection period.

Can the FastAPI WebUI run the crawler without spawning a separate subprocess?

No, according to the MediaCrawler source architecture, the FastAPI WebUI is explicitly designed to delegate to the core crawler via subprocess execution. The manager service in api/services/crawler_manager.py uses asyncio.create_subprocess_exec to launch main.py specifically to isolate the heavy async crawling loop from the web server's event loop, preventing blocking and enabling clean process termination.

Which Pydantic schemas control the communication between the WebUI and crawler?

The integration relies on three primary schemas defined in the API layer: CrawlerStartRequest (validates platform, mode, keywords, and save options on POST), CrawlerStatusResponse (returns running state, PID, and timestamps), and CrawlerLog (structures timestamp, level, and message for log entries). These models ensure type-safe communication between the Vue frontend and the FastAPI control endpoints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →