How QASummaryRestServer Handles Concurrent Request Processing Under Load

QASummaryRestServer uses a dual-layer concurrency model combining FastAPI's asynchronous HTTP handling with Pathway's parallel data-flow engine to process multiple requests simultaneously without blocking.

The QASummaryRestServer class in the pathwaycom/llm-app repository provides REST endpoints for question-answering and summarization workloads. Understanding how this FastAPI-based server handles concurrent request processing is critical for deploying production RAG applications that must serve multiple users under load.

Architecture Overview

The server achieves concurrency through two distinct architectural layers working in tandem. At the network layer, FastAPI and Uvicorn manage asynchronous HTTP connections. At the compute layer, Pathway's streaming runtime parallelizes the actual AI workload execution across CPU cores.

Network-Level Concurrency with FastAPI and Uvicorn

ASGI and Async Coroutines

QASummaryRestServer runs on Uvicorn, an ASGI server that handles HTTP requests as independent coroutines. Located in pathway.xpacks.llm.servers, the server exposes endpoints like /v2/answer and /v2/summarize that operate asynchronously. Because Uvicorn uses an event loop rather than thread-per-connection, receiving request payloads and sending responses never blocks the server, allowing thousands of concurrent connections with minimal memory overhead.

Async Endpoint Handlers

The route handlers inside QASummaryRestServer are defined with async def, enabling them to await the underlying Pathway operations. When a client POSTs to /v2/answer, the handler converts the JSON payload into a query job and yields control back to the event loop while waiting for the result. This coroutine-based approach ensures that a slow LLM inference for one request does not prevent the server from accepting new connections.

Compute-Level Concurrency with Pathway Runtime

Data-Flow Engine

While FastAPI manages the HTTP layer, the actual AI processing occurs within Pathway's runtime (pw.run), which executes a streaming data-flow graph. The runtime is instantiated in templates/question_answering_rag/app.py (lines 63-66) and runs continuously alongside the REST server. This engine is built on a multi-threaded executor that can process many data tokens in parallel across available CPU cores, independent of the Python GIL for its internal operations.

Parallel Query Execution

When QASummaryRestServer receives a request, it injects a new query job into the running Pathway pipeline. The data-flow engine schedules this job on its internal thread pool, allowing multiple queries to undergo embedding generation, vector search, and LLM inference simultaneously. Because the runtime handles its own parallelism, the async Python layer remains unblocked while the heavy compute work scales across hardware resources.

Scaling Strategies for High Load

Multi-Worker Deployment

For deployments requiring horizontal scaling within a single container, Uvicorn supports multiple worker processes via the --workers N flag. Each worker runs an independent Python interpreter with its own QASummaryRestServer instance and Pathway runtime, effectively multiplying throughput on multi-core machines. This approach requires sufficient memory allocation, as each worker maintains its own copy of the model and index.

Container Orchestration

In production environments, QASummaryRestServer is typically deployed behind a load balancer using container platforms like GCP Cloud Run, AWS Fargate, or Azure Container Apps. The combination of FastAPI's async handling and Pathway's parallel engine allows each container instance to maximize its allocated CPU before horizontal scaling kicks in. The stateless design of the REST endpoints ensures that requests can be distributed arbitrarily across a fleet of containers.

Implementation Examples

Starting the Server

The following pattern from templates/question_answering_rag/app.py demonstrates how the server is instantiated alongside the Pathway runtime:


# templates/question_answering_rag/app.py

from pathway.xpacks.llm.servers import QASummaryRestServer
import pathway as pw

class App(BaseModel):
    host: str = "0.0.0.0"
    port: int = 8000
    question_answerer: SummaryQuestionAnswerer

    def run(self) -> None:
        # Initialize the REST server

        server = QASummaryRestServer(self.host, self.port, self.question_answerer)
        
        # Start the Pathway runtime with parallel execution

        pw.run(
            persistence_config=self.persistence_config,
            terminate_on_error=self.terminate_on_error,
            monitoring_level=pw.MonitoringLevel.NONE
        )

Simulating Concurrent Load

Test the server's concurrency handling using Python's ThreadPoolExecutor to fire multiple simultaneous requests:

import requests
from concurrent.futures import ThreadPoolExecutor

ENDPOINT = "http://localhost:8000/v2/answer"
PAYLOAD = {
    "prompt": "What are the key terms of the contract?",
    "filters": {"source": "legal_docs"}
}

def query_server():
    response = requests.post(ENDPOINT, json=PAYLOAD)
    return response.json()

# Execute 50 concurrent requests

with ThreadPoolExecutor(max_workers=50) as executor:
    results = list(executor.map(lambda _: query_server(), range(50)))

print(f"Completed {len(results)} requests")

Multi-Worker Deployment Command

Scale vertically within a single host by launching multiple Uvicorn workers:

uvicorn pathway.xpacks.llm.servers:app \
    --host 0.0.0.0 \
    --port 8000 \
    --workers 4 \
    --loop uvloop

Summary

  • Dual-layer concurrency: QASummaryRestServer combines FastAPI's async HTTP handling with Pathway's parallel data-flow engine to maximize throughput.
  • Non-blocking I/O: Uvicorn's ASGI implementation ensures that network operations never block the event loop, allowing thousands of concurrent connections.
  • Parallel compute: The Pathway runtime distributes query jobs across CPU cores, processing multiple embeddings, retrievals, and LLM inferences simultaneously.
  • Horizontal scaling: The server supports multi-worker processes and containerized deployment for production load balancing.

Frequently Asked Questions

How does QASummaryRestServer prevent request blocking during LLM inference?

The server uses async coroutines in its FastAPI endpoint handlers. When a request requires LLM processing, the handler awaits the Pathway runtime operation, yielding control back to the Uvicorn event loop. This allows the server to accept and process new HTTP connections while previous requests undergo compute-intensive inference in the background.

Can QASummaryRestServer scale across multiple CPU cores on a single machine?

Yes. You can launch the server with multiple Uvicorn workers using the --workers N flag. Each worker runs an independent Python process with its own Pathway runtime instance, effectively multiplying throughput. Alternatively, the Pathway runtime itself parallelizes work across threads within a single process, utilizing all available cores for the data-flow graph execution.

What happens to concurrent requests when the Pathway pipeline is under heavy load?

The Pathway runtime implements backpressure through its data-flow engine. When the pipeline reaches capacity, new query jobs queue briefly in the async layer while the runtime finishes processing current jobs. Because the HTTP layer remains non-blocking, clients maintain their connections, and responses return as soon as the pipeline slots become available. For extreme load, horizontal scaling with container orchestration is recommended.

Is QASummaryRestServer suitable for serverless deployment on platforms like AWS Lambda?

No. QASummaryRestServer is designed for continuous runtime with a persistent Pathway data-flow engine that maintains state (vector indices, embeddings) in memory. Serverless platforms that freeze containers between invocations would destroy this state, causing cold starts that rebuild indices on every request. The server is optimized for containerized deployments with sustained allocation (ECS, Kubernetes, Cloud Run with minimum instances > 0).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →