# How QASummaryRestServer Handles Concurrent Request Processing Under Load

> Discover how QASummaryRestServer efficiently handles concurrent requests under load using FastAPI and Pathway's parallel data-flow for seamless processing.

- Repository: [Pathway/llm-app](https://github.com/pathwaycom/llm-app)
- Tags: performance
- Published: 2026-03-07

---

**`QASummaryRestServer` uses a dual-layer concurrency model combining FastAPI's asynchronous HTTP handling with Pathway's parallel data-flow engine to process multiple requests simultaneously without blocking.**

The `QASummaryRestServer` class in the `pathwaycom/llm-app` repository provides REST endpoints for question-answering and summarization workloads. Understanding how this FastAPI-based server handles concurrent request processing is critical for deploying production RAG applications that must serve multiple users under load.

## Architecture Overview

The server achieves concurrency through two distinct architectural layers working in tandem. At the network layer, FastAPI and Uvicorn manage asynchronous HTTP connections. At the compute layer, Pathway's streaming runtime parallelizes the actual AI workload execution across CPU cores.

## Network-Level Concurrency with FastAPI and Uvicorn

### ASGI and Async Coroutines

`QASummaryRestServer` runs on **Uvicorn**, an ASGI server that handles HTTP requests as independent coroutines. Located in `pathway.xpacks.llm.servers`, the server exposes endpoints like `/v2/answer` and `/v2/summarize` that operate asynchronously. Because Uvicorn uses an event loop rather than thread-per-connection, receiving request payloads and sending responses never blocks the server, allowing thousands of concurrent connections with minimal memory overhead.

### Async Endpoint Handlers

The route handlers inside `QASummaryRestServer` are defined with `async def`, enabling them to **await** the underlying Pathway operations. When a client POSTs to `/v2/answer`, the handler converts the JSON payload into a query job and yields control back to the event loop while waiting for the result. This coroutine-based approach ensures that a slow LLM inference for one request does not prevent the server from accepting new connections.

## Compute-Level Concurrency with Pathway Runtime

### Data-Flow Engine

While FastAPI manages the HTTP layer, the actual AI processing occurs within **Pathway's runtime** (`pw.run`), which executes a streaming data-flow graph. The runtime is instantiated in [`templates/question_answering_rag/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/question_answering_rag/app.py) (lines 63-66) and runs continuously alongside the REST server. This engine is built on a multi-threaded executor that can process many data tokens in parallel across available CPU cores, independent of the Python GIL for its internal operations.

### Parallel Query Execution

When `QASummaryRestServer` receives a request, it injects a new **query job** into the running Pathway pipeline. The data-flow engine schedules this job on its internal thread pool, allowing multiple queries to undergo embedding generation, vector search, and LLM inference simultaneously. Because the runtime handles its own parallelism, the async Python layer remains unblocked while the heavy compute work scales across hardware resources.

## Scaling Strategies for High Load

### Multi-Worker Deployment

For deployments requiring horizontal scaling within a single container, Uvicorn supports multiple worker processes via the `--workers N` flag. Each worker runs an independent Python interpreter with its own `QASummaryRestServer` instance and Pathway runtime, effectively multiplying throughput on multi-core machines. This approach requires sufficient memory allocation, as each worker maintains its own copy of the model and index.

### Container Orchestration

In production environments, `QASummaryRestServer` is typically deployed behind a load balancer using container platforms like GCP Cloud Run, AWS Fargate, or Azure Container Apps. The combination of FastAPI's async handling and Pathway's parallel engine allows each container instance to maximize its allocated CPU before horizontal scaling kicks in. The stateless design of the REST endpoints ensures that requests can be distributed arbitrarily across a fleet of containers.

## Implementation Examples

### Starting the Server

The following pattern from [`templates/question_answering_rag/app.py`](https://github.com/pathwaycom/llm-app/blob/main/templates/question_answering_rag/app.py) demonstrates how the server is instantiated alongside the Pathway runtime:

```python

# templates/question_answering_rag/app.py

from pathway.xpacks.llm.servers import QASummaryRestServer
import pathway as pw

class App(BaseModel):
    host: str = "0.0.0.0"
    port: int = 8000
    question_answerer: SummaryQuestionAnswerer

    def run(self) -> None:
        # Initialize the REST server

        server = QASummaryRestServer(self.host, self.port, self.question_answerer)
        
        # Start the Pathway runtime with parallel execution

        pw.run(
            persistence_config=self.persistence_config,
            terminate_on_error=self.terminate_on_error,
            monitoring_level=pw.MonitoringLevel.NONE
        )

```

### Simulating Concurrent Load

Test the server's concurrency handling using Python's `ThreadPoolExecutor` to fire multiple simultaneous requests:

```python
import requests
from concurrent.futures import ThreadPoolExecutor

ENDPOINT = "http://localhost:8000/v2/answer"
PAYLOAD = {
    "prompt": "What are the key terms of the contract?",
    "filters": {"source": "legal_docs"}
}

def query_server():
    response = requests.post(ENDPOINT, json=PAYLOAD)
    return response.json()

# Execute 50 concurrent requests

with ThreadPoolExecutor(max_workers=50) as executor:
    results = list(executor.map(lambda _: query_server(), range(50)))

print(f"Completed {len(results)} requests")

```

### Multi-Worker Deployment Command

Scale vertically within a single host by launching multiple Uvicorn workers:

```bash
uvicorn pathway.xpacks.llm.servers:app \
    --host 0.0.0.0 \
    --port 8000 \
    --workers 4 \
    --loop uvloop

```

## Summary

- **Dual-layer concurrency**: `QASummaryRestServer` combines FastAPI's async HTTP handling with Pathway's parallel data-flow engine to maximize throughput.
- **Non-blocking I/O**: Uvicorn's ASGI implementation ensures that network operations never block the event loop, allowing thousands of concurrent connections.
- **Parallel compute**: The Pathway runtime distributes query jobs across CPU cores, processing multiple embeddings, retrievals, and LLM inferences simultaneously.
- **Horizontal scaling**: The server supports multi-worker processes and containerized deployment for production load balancing.

## Frequently Asked Questions

### How does QASummaryRestServer prevent request blocking during LLM inference?

The server uses **async coroutines** in its FastAPI endpoint handlers. When a request requires LLM processing, the handler awaits the Pathway runtime operation, yielding control back to the Uvicorn event loop. This allows the server to accept and process new HTTP connections while previous requests undergo compute-intensive inference in the background.

### Can QASummaryRestServer scale across multiple CPU cores on a single machine?

Yes. You can launch the server with **multiple Uvicorn workers** using the `--workers N` flag. Each worker runs an independent Python process with its own Pathway runtime instance, effectively multiplying throughput. Alternatively, the Pathway runtime itself parallelizes work across threads within a single process, utilizing all available cores for the data-flow graph execution.

### What happens to concurrent requests when the Pathway pipeline is under heavy load?

The Pathway runtime implements **backpressure** through its data-flow engine. When the pipeline reaches capacity, new query jobs queue briefly in the async layer while the runtime finishes processing current jobs. Because the HTTP layer remains non-blocking, clients maintain their connections, and responses return as soon as the pipeline slots become available. For extreme load, horizontal scaling with container orchestration is recommended.

### Is QASummaryRestServer suitable for serverless deployment on platforms like AWS Lambda?

No. `QASummaryRestServer` is designed for **continuous runtime** with a persistent Pathway data-flow engine that maintains state (vector indices, embeddings) in memory. Serverless platforms that freeze containers between invocations would destroy this state, causing cold starts that rebuild indices on every request. The server is optimized for containerized deployments with sustained allocation (ECS, Kubernetes, Cloud Run with minimum instances > 0).