How to Configure Concurrency Limits and Rate Limiting in Prefect Work Pools
To configure concurrency limits in Prefect work pools, set the concurrency_limit field on the WorkPool or WorkQueue models via the CLI or API, and implement rate limiting using global concurrency limits with the rate_limit helper in prefect.concurrency.sync or asyncio.
Prefect 3 uses work pools to route flow runs to workers, allowing you to control execution capacity at both the infrastructure and task levels. Whether you need to cap simultaneous runs in a Kubernetes cluster or throttle API calls to external services, understanding how to configure concurrency limits and rate limiting in Prefect work pools is essential for production orchestration. This guide references the actual implementation in the PrefectHQ/prefect repository to show you exactly how these mechanisms work.
Understanding Concurrency Limits vs. Rate Limiting
Prefect provides two distinct throttling mechanisms that operate at different layers of the orchestration stack:
-
Concurrency limits control the maximum number of flow runs that can execute simultaneously within a work pool or specific work queue. These are stored in the
concurrency_limitfield of theWorkPoolandWorkQueuemodels defined insrc/prefect/server/models/workers.pyandsrc/prefect/server/models/work_queues.py. -
Rate limiting controls the frequency at which operations may start, implemented via global concurrency limits where the
slot_decay_per_secondparameter defines the allowed rate. This logic lives insrc/prefect/concurrency/sync.py(line 83) andsrc/prefect/concurrency/asyncio.py(line 86).
The server enforces pool and queue concurrency limits during the worker polling cycle, while rate limits are enforced client-side using the SDK helpers.
Configuring Work Pool Concurrency Limits
Work pool limits define the total capacity for all queues within that pool. When set, the server counts active flow runs (states RUNNING and PENDING) against this ceiling before allocating new work.
Setting Limits via CLI
Use the set-concurrency-limit command to impose a pool-wide cap:
# Limit the "default" pool to 10 simultaneous flow runs
prefect work-pool set-concurrency-limit default --limit 10
This command updates the WorkPool.concurrency_limit field in the database and displays a visual slots indicator. The implementation resides in src/prefect/cli/work_pool.py (lines 794-825), which calls the server API to persist the change.
Server-Side Enforcement
When a worker polls for work, the server queries the active run count for the target pool. If the count meets or exceeds concurrency_limit, the API returns no runs, causing the worker to pause. The relevant counting logic appears in src/prefect/server/models/workers.py, where the orchestration layer filters available work based on the current occupancy.
Configuring Work Queue Concurrency Limits
Individual work queues can impose stricter limits than their parent pool, enabling priority-based capacity management.
Queue-Level Overrides
The WorkQueue model (defined in src/prefect/server/models/work_queues.py) contains its own concurrency_limit field. Queues inherit the pool limit by default, but setting a specific value on the queue creates a stricter ceiling. The server calculates open_concurrency_slots in lines 578-589 of work_queues.py, returning the lesser of the pool or queue availability.
CLI Commands
Create a queue with an initial limit or update an existing one:
# Create a high-priority queue limited to 3 concurrent runs
prefect work-queue create high-priority --priority 5 --concurrency-limit 3
# Update an existing queue
prefect work-queue set-concurrency-limit high-priority --limit 3
These commands are implemented in src/prefect/cli/work_queue.py (lines 137-165). If the limit is reached, new flow runs remain in the awaiting concurrency slot state until a slot becomes available.
Implementing Rate Limiting with Global Concurrency Limits
For throttling access to external APIs or shared resources, use global concurrency limits with the rate_limit helper.
Creating Global Concurrency Limits
A global concurrency limit is a first-class object that can be created via the UI or API. When its mode is set to "rate_limit", the slot_decay_per_second field determines how quickly occupied slots become available again. For example, a decay value of 0.2 permits 5 calls per second.
Using the rate_limit Helper (Sync)
The rate_limit context manager in src/prefect/concurrency/sync.py (line 83) acquires a slot and blocks until one is available:
from prefect.concurrency.sync import rate_limit
def call_third_party_api():
# Acquire a slot; blocks if the rate limit is exceeded
rate_limit("api-rate", occupy=1)
# ... perform API request ...
# Slot automatically releases after the decay interval
Using the rate_limit Helper (Async)
For asynchronous workflows, use the equivalent in src/prefect/concurrency/asyncio.py (line 86):
from prefect.concurrency.asyncio import rate_limit
async def call_third_party_api():
await rate_limit("api-rate", occupy=1)
# ... async API request ...
The SDK automatically creates the underlying global concurrency limit if it does not exist, handling the negotiation with the server via src/prefect/client/orchestration/_concurrency_limits/client.py.
How the Server Enforces Limits
The enforcement mechanism relies on SQL queries that count running runs against the configured limits. In src/prefect/server/models/work_queues.py, the work_queue_concurrency_slots() function checks if work_queue.concurrency_limit is not None: (line 578) and caps the returned slots accordingly. Similarly, work pool limits are validated in src/prefect/server/models/workers.py before the API returns work to polling workers.
This architecture ensures that workers remain simple polling clients while the server maintains centralized control over capacity and frequency constraints.
Summary
- Concurrency limits cap simultaneous execution and are configured via
WorkPool.concurrency_limitandWorkQueue.concurrency_limitin the server models. - Rate limits control execution frequency using global concurrency limits with
slot_decay_per_secondand therate_limithelper functions. - Use the CLI commands
prefect work-pool set-concurrency-limitandprefect work-queue set-concurrency-limitto configure capacity limits quickly. - The server enforces limits by counting active runs in
RUNNINGorPENDINGstates before returning work to polling workers. - Global rate limits work across both sync and async code via
prefect.concurrency.syncandprefect.concurrency.asyncio.
Frequently Asked Questions
What is the difference between a work pool concurrency limit and a work queue concurrency limit?
A work pool concurrency limit sets the total capacity for all queues within that pool, defined in the WorkPool model in src/prefect/server/models/workers.py. A work queue concurrency limit provides finer-grained control for specific priority lanes within the pool, stored in the WorkQueue model in src/prefect/server/models/work_queues.py. The server checks both limits and applies the more restrictive value when determining whether to allocate a run.
How does Prefect handle rate limiting internally?
Prefect implements rate limiting through global concurrency limits that use a slot-based decay mechanism. When you call rate_limit(), the SDK acquires a slot from the limit object. If the limit's mode is "rate_limit", slots become available again at the rate specified by slot_decay_per_second. The implementation in src/prefect/concurrency/sync.py (line 83) and asyncio.py (line 86) blocks execution until a slot is available, effectively throttling the operation frequency.
Can I set concurrency limits programmatically instead of using the CLI?
Yes. While the CLI provides convenience commands in src/prefect/cli/work_pool.py and work_queue.py, you can set limits programmatically by updating the concurrency_limit field on WorkPool or WorkQueue objects via the Prefect API client. The Python SDK also allows creating and managing global concurrency limits for rate limiting scenarios through the orchestration client at src/prefect/client/orchestration/_concurrency_limits/client.py.
What happens to flow runs when the concurrency limit is reached?
When a work pool or queue reaches its concurrency_limit, the server stops allocating new runs to workers polling that pool or queue. The flow runs remain in an awaiting concurrency slot state (as documented in the states concept documentation) until an active run completes and frees up capacity. Workers continue polling normally, and the server automatically assigns work once slots become available without requiring manual intervention.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →