Configuring Concurrent Crawling with Per-Domain Throttling in Scrapling
Scrapling enables fine-grained concurrency control through the Spider class attributes concurrent_requests and concurrent_requests_per_domain, enforced by CapacityLimiter objects in the CrawlerEngine.
Configuring concurrent crawling with per-domain throttling in Scrapling allows you to maximize throughput while respecting server rate limits. The D4Vinci/Scrapling repository implements this through a three-layer architecture involving the Spider configuration class, the CrawlerEngine execution orchestrator, and the Scheduler queue manager.
Core Components of Scrapling's Concurrency Model
Scrapling's crawling engine relies on three interconnected components to manage concurrent execution:
Spiderclass – Defined inscrapling/spiders/spider.py, this base class holds the concurrency configuration as class attributes.CrawlerEngine– Located inscrapling/spiders/engine.py, this class orchestrates request execution using AnyIOCapacityLimiterobjects to enforce concurrency rules.Scheduler– Implemented inscrapling/spiders/scheduler.py, this priority queue feeds the engine with deduplicatedRequestobjects.
Where Concurrency Settings Are Defined
The Spider class in scrapling/spiders/spider.py declares three critical attributes that control crawling behavior:
| Setting | Default | Description |
|---|---|---|
concurrent_requests |
4 |
Global maximum number of simultaneous requests across all domains. |
concurrent_requests_per_domain |
0 |
Maximum concurrent requests allowed per individual domain (0 disables per-domain limiting). |
download_delay |
0.0 |
Minimum delay in seconds between requests to the same domain. |
When concurrent_requests_per_domain is set to a value greater than 0, the CrawlerEngine activates per-domain throttling by creating dedicated CapacityLimiter instances for each domain encountered during the crawl.
How Per-Domain Throttling Works
The CrawlerEngine in scrapling/spiders/engine.py implements throttling through the _rate_limiter method and the _process_request coroutine.
When processing a request, the engine wraps the download call in an async context manager:
async with self._rate_limiter(request.domain):
if self.spider.download_delay:
await anyio.sleep(self.spider.download_delay)
response = await self.session_manager.fetch(request)
The _rate_limiter method determines which limiter to apply:
- Global limiting only – If
concurrent_requests_per_domainis0, all requests share the globalCapacityLimiterinitialized withconcurrent_requests. - Per-domain limiting – If enabled, the first request to a domain creates a dedicated
CapacityLimiterstored inself._domain_limiters[domain]. Subsequent requests to that domain acquire slots from this specific limiter.
This architecture ensures that at most concurrent_requests_per_domain requests run concurrently against any single domain, while the global concurrent_requests cap prevents total resource exhaustion across all domains.
Practical Implementation Example
The following example demonstrates configuring a spider with both global and per-domain concurrency limits:
# example_spider.py
from scrapling.spiders.spider import Spider, Request
class MySpider(Spider):
name = "my_spider"
start_urls = [
"https://example.com",
"https://blog.example.com",
"https://other.com",
]
# Global cap – at most 8 parallel requests overall
concurrent_requests = 8
# Per‑domain cap – at most 2 concurrent requests per domain
concurrent_requests_per_domain = 2
download_delay = 0.2 # 200 ms between requests to the same domain
async def parse(self, response):
# Your parsing logic here – you can yield new Request objects
# They will automatically respect the limits defined above.
yield {
"url": str(response.url),
"status": response.status,
"length": len(response.body),
}
# Running the spider
if __name__ == "__main__":
from scrapling.spiders.engine import CrawlerEngine
from scrapling.spiders.session import SessionManager
spider = MySpider()
engine = CrawlerEngine(spider, SessionManager())
# Simple async entry point
import anyio
anyio.run(engine.crawl)
In this configuration:
concurrent_requests = 8creates a globalCapacityLimiter(8)that restricts the engine to eight simultaneous requests across all domains.concurrent_requests_per_domain = 2activates per-domain limiters, ensuring thatexample.com,blog.example.com, andother.comeach process at most two requests concurrently.download_delay = 0.2inserts a 200-millisecond pause after acquiring a limiter slot, further reducing load on target servers.
Summary
- Global concurrency is controlled by the
concurrent_requestsattribute inscrapling/spiders/spider.py, enforced by a globalCapacityLimiterinCrawlerEngine. - Per-domain throttling activates when
concurrent_requests_per_domainis greater than0, creating dedicatedCapacityLimiterinstances for each domain inscrapling/spiders/engine.py. - Download delays are applied after acquiring a limiter slot via
anyio.sleep(), ensuring respectful crawling behavior. - The
Schedulermanages request deduplication and priority, feeding the engine while concurrency controls regulate execution speed.
Frequently Asked Questions
What is the default concurrency limit in Scrapling?
By default, Scrapling allows four concurrent requests globally (concurrent_requests = 4), while per-domain throttling is disabled (concurrent_requests_per_domain = 0). These defaults are defined in the Spider base class located at scrapling/spiders/spider.py.
How does Scrapling handle per-domain throttling when concurrent_requests_per_domain is set to 0?
When concurrent_requests_per_domain is 0, Scrapling disables per-domain limiting entirely. All requests share the global CapacityLimiter created in CrawlerEngine.__init__. This means a single domain could potentially consume all available global slots, though the Scheduler and engine task management prevent total resource exhaustion.
Can I set different download delays for different domains?
Currently, Scrapling implements a global download_delay attribute on the Spider class that applies uniformly to all domains. The delay is enforced in CrawlerEngine._process_request after acquiring the rate limiter slot. To implement domain-specific delays, you would need to subclass CrawlerEngine and override the _process_request method to inspect request.domain before applying anyio.sleep().
What happens if the global concurrent_requests limit is lower than concurrent_requests_per_domain?
If concurrent_requests is set lower than concurrent_requests_per_domain, the global limit takes precedence. The CrawlerEngine creates a global CapacityLimiter with the concurrent_requests value, and all requests—regardless of domain—must acquire a slot from this global limiter before the per-domain limiter is even consulted. This ensures that the total number of in-flight requests never exceeds the global cap, protecting system resources.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →