# How the FastAPI WebUI Integrates with the MediaCrawler Core Engine

> Explore how the FastAPI WebUI integrates with the MediaCrawler core engine. This guide details its role in managing the crawler's lifecycle, logs, and status through REST and WebSockets.

- Repository: [程序员阿江-Relakkes/MediaCrawler](https://github.com/NanmiCoder/MediaCrawler)
- Tags: architecture
- Published: 2026-07-02

---

**The FastAPI WebUI in MediaCrawler acts as a thin HTTP control plane that manages the crawling engine through a dedicated service layer, handling process lifecycle, log streaming, and status monitoring via REST endpoints and WebSockets.**

The MediaCrawler open-source project provides a browser-based interface for controlling multi-platform content crawlers. Rather than embedding the crawling logic directly into the web server, the FastAPI WebUI operates as a separate orchestration layer that spawns and monitors the core crawler as an asynchronous subprocess, enabling safe concurrent operation and real-time log streaming.

## FastAPI Application Structure

### Main Application Setup

The FastAPI application is initialized in [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py), where the server registers CORS middleware and mounts the compiled Vue/Vite frontend bundle. Static assets are served under `/assets`, `/logos`, and `/static`, while the root endpoint (`GET /`) returns [`index.html`](https://github.com/NanmiCoder/MediaCrawler/blob/main/index.html) to bootstrap the single-page application.

```python

# From api/main.py - FastAPI instantiation and static mounting

app = FastAPI()
app.mount("/assets", StaticFiles(directory="api/webui/assets"), name="assets")

```

### Router Architecture

The API routes are organized into three specialized routers defined in [`api/routers/__init__.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/__init__.py):

- **`crawler_router`** – Controls crawler lifecycle (`/api/crawler/start`, `/api/crawler/stop`, `/api/crawler/status`, `/api/crawler/logs`)
- **`data_router`** – Handles data export operations (JSON, CSV, Excel downloads)
- **`websocket_router`** – Provides real-time log streaming capabilities

These routers are included with a common `/api` prefix, creating a clean separation between the web interface and the underlying crawling engine.

## Crawler Control Endpoints

The `crawler_router` (implemented in [`api/routers/crawler.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/routers/crawler.py)) delegates all operations to the **crawler manager service** rather than executing crawl logic directly. This design prevents the FastAPI event loop from blocking during long-running scraping operations.

### Starting the Crawler

When the frontend sends a `POST` request to `/api/crawler/start`, the endpoint receives a `CrawlerStartRequest` Pydantic model containing platform selection, crawling mode, keywords, and authentication options. The router passes this to `crawler_manager.start(request)`, which launches the actual crawler entry point ([`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py)) as an async subprocess using `asyncio.create_subprocess_exec`.

```python

# Conceptual flow from api/routers/crawler.py

@crawler_router.post("/start")
async def start_crawler(request: CrawlerStartRequest):
    return await crawler_manager.start(request)

```

### Stopping and Monitoring

The control endpoints provide safe process management:

- **`POST /api/crawler/stop`** – Invokes `crawler_manager.stop()`, which terminates the subprocess cleanly using `process.terminate()` and monitors the exit code
- **`GET /api/crawler/status`** – Returns a `CrawlerStatusResponse` containing the running state, process ID (`pid`), start timestamp, and current log line count by checking `process.poll()`
- **`GET /api/crawler/logs?limit=n`** – Retrieves the most recent entries from the manager's circular log buffer

```javascript
// Polling crawler status from the frontend
setInterval(async () => {
  const resp = await fetch('/api/crawler/status');
  const status = await resp.json();
  console.log('Crawler running:', status.running, 'PID:', status.pid);
}, 2000);

```

## The Crawler Manager Service

The [`crawler_manager.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/crawler_manager.py) service encapsulates all interaction with the crawling engine, maintaining singleton process state and providing concurrency safety.

### Process Management

The service maintains an internal `process` reference that tracks the running crawler subprocess. Before starting a new crawl, it checks `process.poll()` to ensure no existing process is running, returning HTTP 400 if a duplicate start request is detected. The manager constructs the command-line invocation using `uv run main.py` with the parameters extracted from the `CrawlerStartRequest` model.

### Log Collection and Status Tracking

While the crawler runs, the manager asynchronously reads from subprocess stdout and stderr, parsing each line into a `CrawlerLog` Pydantic model with timestamp, log level, and message fields. These entries are stored in a bounded list for quick retrieval by the logs endpoint. The `CrawlerStatusResponse` schema exposed by the service provides the frontend with accurate runtime metadata including `running`, `pid`, `started_at`, and `log_line_count`.

```python

# From api/services/crawler_manager.py

class CrawlerStatusResponse(BaseModel):
    running: bool
    pid: int | None
    started_at: datetime | None
    log_line_count: int

```

## Frontend Integration

The Vue-based frontend (bundled in `api/webui`) communicates with the FastAPI layer through standard HTTP requests and WebSocket connections. When initiating a crawl, the UI serializes the configuration into a `CrawlerStartRequest` JSON payload:

```javascript
await fetch('/api/crawler/start', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    platform: 'dy',
    mode: 'search',
    keyword: 'funny cats',
    save_option: 'json',
    login_type: 'qrcode'
  })
});

```

For real-time updates, the frontend either polls `/api/crawler/status` and `/api/crawler/logs` periodically or connects to the WebSocket endpoint for push-based log streaming. This architecture decouples the heavy I/O of web scraping from the web server's event loop while keeping both components within the same Python process space for easy deployment.

## Summary

- **FastAPI WebUI** serves as a control plane in [`api/main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/main.py), mounting static frontend assets and registering route prefixes
- **Three dedicated routers** (`crawler_router`, `data_router`, `websocket_router`) handle distinct concerns under the `/api` prefix
- **Crawler manager service** ([`api/services/crawler_manager.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/services/crawler_manager.py)) encapsulates subprocess lifecycle, using `asyncio.create_subprocess_exec` to launch [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) and monitoring via `process.poll()`
- **Pydantic models** (`CrawlerStartRequest`, `CrawlerStatusResponse`, `CrawlerLog`) enforce type safety between the UI and crawling engine
- **Real-time log streaming** is achieved through asynchronous stdout/stderr reading into a bounded buffer, accessible via REST or WebSocket endpoints

## Frequently Asked Questions

### How does the FastAPI WebUI prevent starting multiple crawler instances simultaneously?

The crawler manager service maintains a singleton `process` reference and checks `process.poll()` before accepting new start requests. If a process is already running, it returns HTTP 400 with a clear error message indicating that a crawl is currently active, ensuring safe concurrency control at the application layer.

### What happens to crawler logs when the WebUI is not actively polling?

The crawler manager stores log entries in an in-memory circular buffer (bounded list) while the subprocess runs. Even if the frontend disconnects or stops polling, logs continue accumulating in this buffer up to the configured limit. When the client resumes polling `/api/crawler/logs`, it retrieves the most recent entries from this buffer, including those generated during the disconnection period.

### Can the FastAPI WebUI run the crawler without spawning a separate subprocess?

No, according to the MediaCrawler source architecture, the FastAPI WebUI is explicitly designed to delegate to the core crawler via subprocess execution. The manager service in [`api/services/crawler_manager.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/api/services/crawler_manager.py) uses `asyncio.create_subprocess_exec` to launch [`main.py`](https://github.com/NanmiCoder/MediaCrawler/blob/main/main.py) specifically to isolate the heavy async crawling loop from the web server's event loop, preventing blocking and enabling clean process termination.

### Which Pydantic schemas control the communication between the WebUI and crawler?

The integration relies on three primary schemas defined in the API layer: `CrawlerStartRequest` (validates platform, mode, keywords, and save options on POST), `CrawlerStatusResponse` (returns running state, PID, and timestamps), and `CrawlerLog` (structures timestamp, level, and message for log entries). These models ensure type-safe communication between the Vue frontend and the FastAPI control endpoints.