# How to Use Streaming Mode for Real-Time Crawling Stats and Item Processing in Scrapling

> Learn to use Scrapling's streaming mode for real-time crawling stats and item processing. Access live CrawlStats and scraped items immediately with the Spider.stream() method.

- Repository: [Karim shoair/Scrapling](https://github.com/D4Vinci/Scrapling)
- Tags: how-to-guide
- Published: 2026-03-08

---

**Scrapling’s `Spider.stream()` method provides an async generator that yields scraped items immediately upon extraction while exposing live `CrawlStats` for real-time monitoring, implemented via a bounded in-memory channel in `CrawlerEngine`.**

Scrapling is an open-source asynchronous web crawling framework designed for high-performance data extraction. When you need to process scraped items as they arrive rather than waiting for an entire crawl to complete, Scrapling’s streaming mode offers a robust solution for real-time crawling stats and item processing without buffering everything in memory.

## Understanding Scrapling’s Streaming Architecture

Scrapling’s streaming capability relies on a tight integration between the `Spider` class and the `CrawlerEngine`, using async generators and memory object streams to decouple production from consumption.

### Core Components

| Component | Role | Key Implementation |
|-----------|------|--------------------|
| **`Spider.stream`** | Public async generator that yields items one-by-one and gives access to `spider.stats` during iteration. | [`scrapling/spiders/spider.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/spider.py) – lines 90-103 |
| **`CrawlerEngine.__aiter__` / `_stream`** | The engine drives the crawl, pushes each processed result into an internal memory object stream, and yields it to the spider. | [`scrapling/spiders/engine.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/engine.py) – lines 13-34 |
| **`create_memory_object_stream`** | Provides a bounded async queue (size 100) that decouples crawling tasks from the consumer, enabling back-pressure-aware streaming. | Imported from **anyio** in [`engine.py`](https://github.com/D4Vinci/Scrapling/blob/main/engine.py) |
| **`Spider.stats`** | Proxy property that forwards to `engine.stats` while a stream is active, allowing real-time inspection of the `CrawlStats` dataclass. | [`scrapling/spiders/spider.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/spider.py) – lines 11-16 |
| **`CrawlStats`** | Holds counters (requests, items scraped, bytes, response codes, etc.) that are updated continuously during the crawl. | [`scrapling/spiders/result.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/result.py) – lines 41-67 |

### Data Flow

1. **`spider.stream()`** creates a logger token, instantiates `CrawlerEngine`, and enters `async for item in self._engine`.
2. The engine spawns a **task group**; `run()` performs the full crawl (`await self.crawl()`).
3. Inside `_stream`, `create_memory_object_stream` returns a **sender (`send`)** and **receiver (`recv`)** pair. Every processed result is pushed via `await self._item_stream.send(processed_result)` (see [`engine.py`](https://github.com/D4Vinci/Scrapling/blob/main/engine.py) lines 144-146).
4. The spider’s `async for` loop consumes items from `recv`, yielding them to the caller in real time.
5. While the loop runs, `spider.stats` reads the live `engine.stats` object, which is continuously updated by methods such as `increment_requests_count`, `increment_status`, etc., defined in `CrawlStats`.

Because the stream uses an **in-memory bounded channel**, the crawl will pause when the consumer falls behind, preventing unbounded memory growth.

## Accessing Real-Time Crawling Statistics

The `CrawlStats` dataclass tracks comprehensive metrics that update continuously during streaming:

- `requests_count`, `failed_requests_count`, `blocked_requests_count`
- `items_scraped`, `items_dropped`
- `response_status_count` (HTTP-code histogram)
- `response_bytes` and per-domain byte counters
- `download_delay`, `concurrent_requests`, and `concurrent_requests_per_domain`
- Custom metrics (`custom_stats`) and logger level counters (`log_levels_counter`)

All fields are updated during crawling; the `spider.stats` property makes them instantly readable inside the streaming loop.

## Practical Implementation Examples

### Basic Streaming Loop

```python
import asyncio
from scrapling.spiders.spider import Spider

class MySpider(Spider):
    start_urls = ["https://example.com"]

    async def parse(self, response):
        # Yield a dict for each page

        yield {"url": response.url, "title": response.title}
        # Optionally generate follow-up requests

        if "next" in response.links:
            yield Request(response.links["next"], sid="default")

async def main():
    spider = MySpider()
    async for item in spider.stream():
        print("Item:", item)
        # Live stats are always available

        print("Requests so far:", spider.stats.requests_count)

asyncio.run(main())

```

**Key points:**

- `spider.stream()` must be called from an **async context**.
- Inside the loop, `spider.stats` reflects the latest crawl status.

### Live Dashboard Integration

```python
import asyncio
import websockets
from scrapling.spiders.spider import Spider

class DashboardSpider(Spider):
    start_urls = ["https://news.ycombinator.com"]

    async def parse(self, response):
        yield {"title": response.title, "url": response.url}

async def feed_websocket():
    async with websockets.connect("ws://localhost:8765") as ws:
        spider = DashboardSpider()
        async for item in spider.stream():
            await ws.send(json.dumps({"item": item, "stats": spider.stats.__dict__}))

asyncio.run(feed_websocket())

```

Each item is pushed immediately to a WebSocket client, together with a snapshot of the current `CrawlStats`.

### Processing Follow-Up Requests in Stream Mode

```python
async def parse(response):
    # Yield a new request before emitting the current item

    if "details" in response.links:
        yield Request(response.links["details"], sid="default")
    yield {"url": response.url, "summary": response.text[:200]}

spider = MySpider()
spider.parse = parse   # Swap parser at runtime

async for item in spider.stream():
    print(item)   # Will include items from the follow-up request as they arrive

```

The test suite ([`tests/spiders/test_engine.py`](https://github.com/D4Vinci/Scrapling/blob/main/tests/spiders/test_engine.py)) validates exactly this behavior (see lines 22-30 for the follow-up request test).

## Key Source Files and Implementation Details

| File | Purpose | Link |
|------|---------|------|
| [`scrapling/spiders/spider.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/spider.py) | Public spider API, `stream()` generator, `stats` proxy | [spider.py](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/spider.py) |
| [`scrapling/spiders/engine.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/engine.py) | Core crawling engine, async iterator, memory stream handling | [engine.py](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/engine.py) |
| [`scrapling/spiders/result.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/result.py) | `CrawlStats` dataclass and `CrawlResult` wrapper | [result.py](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/result.py) |
| [`tests/spiders/test_engine.py`](https://github.com/D4Vinci/Scrapling/blob/main/tests/spiders/test_engine.py) | Test suite covering streaming, stats, and pause logic | [test_engine.py](https://github.com/D4Vinci/Scrapling/blob/main/tests/spiders/test_engine.py) |
| [`scrapling/cli.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/cli.py) | CLI entry point – the `--stream` flag can trigger streaming mode from the command line | [cli.py](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/cli.py) |

These files together implement the streaming capability, provide real-time statistics, and expose a clean async API for developers.

## Summary

- **Streaming mode** in Scrapling is enabled via `spider.stream()`, an async generator that yields items as they are scraped rather than batching results at the end.
- The architecture uses `create_memory_object_stream` from **anyio** to create a bounded queue (size 100) that decouples the crawler from the consumer, providing automatic back-pressure handling.
- Real-time statistics are available through the `spider.stats` property, which proxies to `engine.stats` and exposes the `CrawlStats` dataclass with counters for requests, items, bytes, and HTTP status codes.
- The implementation spans [`scrapling/spiders/spider.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/spider.py) for the public API, [`scrapling/spiders/engine.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/engine.py) for the async iteration logic, and [`scrapling/spiders/result.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/result.py) for the statistics dataclass.

## Frequently Asked Questions

### How does Scrapling handle backpressure in streaming mode?

Scrapling uses a bounded memory object stream with a default buffer size of 100 items, implemented via `create_memory_object_stream` from the **anyio** library in [`scrapling/spiders/engine.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/engine.py). When the consumer falls behind, the crawler automatically pauses until the buffer drains, preventing unbounded memory growth during long-running crawls.

### Can I access crawl statistics outside of the streaming loop?

The `spider.stats` property is specifically designed to be accessed during active streaming, as it proxies to the live `engine.stats` object maintained by `CrawlerEngine`. While the stats object exists after crawling completes, the real-time updates only occur during the `async for` iteration over `spider.stream()`, making the streaming loop the intended context for monitoring `CrawlStats`.

### What types of objects can be yielded when using streaming mode?

Within your `parse` method, you can yield dictionaries representing scraped items or `Request` objects to schedule follow-up crawls. The `CrawlerEngine` in [`scrapling/spiders/engine.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/spiders/engine.py) processes both types through its `_stream` method, pushing them into the memory object stream where they are yielded to your consumer loop in real-time, as validated in [`tests/spiders/test_engine.py`](https://github.com/D4Vinci/Scrapling/blob/main/tests/spiders/test_engine.py) lines 22-30.

### How do I enable streaming mode from the command line?

Scrapling’s CLI interface in [`scrapling/cli.py`](https://github.com/D4Vinci/Scrapling/blob/main/scrapling/cli.py) supports a `--stream` flag that triggers streaming mode when invoking spiders from the terminal. When this flag is active, the CLI consumes the async generator returned by `spider.stream()` and outputs items as they are scraped, rather than collecting them into a final batch for post-crawl processing.