# Understanding Crawl4AI Caching Modes: BYPASS, READ_ONLY, WRITE_ONLY, and More

> Explore Crawl4AI caching modes like BYPASS, READ_ONLY, and WRITE_ONLY. Understand how these options control disk access for efficient web crawling and data management with the AsyncWebCrawler.

- Repository: [UncleCode/crawl4ai](https://github.com/unclecode/crawl4ai)
- Tags: deep-dive
- Published: 2026-03-05

---

**The `CacheMode` enumeration in crawl4ai provides five explicit options—`ENABLED`, `DISABLED`, `READ_ONLY`, `WRITE_ONLY`, and `BYPASS`—to control how the `AsyncWebCrawler` reads from or writes to disk, replacing legacy boolean flags with type-safe, self-documenting behavior.**

Modern web crawling pipelines need precise control over disk I/O to balance performance against data freshness. The crawl4ai library (maintained at `unclecode/crawl4ai`) centralizes this decision in a single `CacheMode` enum defined in [`crawl4ai/cache_context.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/cache_context.py), making it trivial to enforce read-only validation runs, write-only cache warm-ups, or completely bypass persistent storage.

## The Five Crawl4AI Caching Modes Explained

The `CacheMode` enum defines five distinct behaviors. Each mode explicitly declares whether the crawler may read cached content and whether it may write new content to the cache directory.

### CacheMode.ENABLED

**`CacheMode.ENABLED`** is the default behavior. The crawler reads from the cache when a fresh entry exists and writes newly fetched pages back to disk for future reuse. Use this for standard production workloads where you want automatic deduplication and offline replay capability.

### CacheMode.DISABLED

**`CacheMode.DISABLED`** disables caching entirely. The crawler never checks for existing files and never writes results to disk, eliminating all cache-related I/O. Choose this mode for one-off extractions where you want to avoid polluting the filesystem or when running in memory-constrained environments.

### CacheMode.READ_ONLY

**`CacheMode.READ_ONLY`** allows the crawler to fetch from existing cache entries but strictly forbids writing new data. This is ideal for validation pipelines or CI jobs where you must guarantee that the crawler never mutates the cache directory, ensuring complete reproducibility from existing snapshots.

### CacheMode.WRITE_ONLY

**`CacheMode.WRITE_ONLY`** forces the crawler to ignore existing cache entries while still persisting every fetched page to disk. Use this for "warm-up" runs that populate a shared cache without being affected by stale or incomplete historical data.

### CacheMode.BYPASS

**`CacheMode.BYPASS`** completely skips the cache for the current operation—neither reading nor writing. This differs subtly from `DISABLED`: while functionally similar for a single run, `BYPASS` is intended for scenarios where you temporarily need fresh data (e.g., real-time extraction or debugging) while keeping the global cache intact for other operations.

## How CacheMode Works Under the Hood

The enum definition resides in [`crawl4ai/cache_context.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/cache_context.py) at lines 4-21:

```python
class CacheMode(Enum):
    """
    Defines the caching behavior for web crawling operations.
    """
    ENABLED = "enabled"
    DISABLED = "disabled"
    READ_ONLY = "read_only"
    WRITE_ONLY = "write_only"
    BYPASS = "bypass"

```

When a crawl starts, the library instantiates a `CacheContext` object that translates the selected mode into boolean decisions via the `should_read()` and `should_write()` methods (lines 59-88 in the same file):

```python
def should_read(self) -> bool:
    if self.always_bypass or not self.is_cacheable:
        return False
    return self.cache_mode in [CacheMode.ENABLED, CacheMode.READ_ONLY]

def should_write(self) -> bool:
    if self.always_bypass or not self.is_cacheable:
        return False
    return self.cache_mode in [CacheMode.ENABLED, CacheMode.WRITE_ONLY]

```

These methods demonstrate the internal logic: `READ_ONLY` appears only in the read whitelist, `WRITE_ONLY` appears only in the write whitelist, and `ENABLED` appears in both. The `BYPASS` and `DISABLED` modes are excluded from both lists, though `BYPASS` additionally triggers the `always_bypass` guard clause.

## Practical Usage Examples

Configure caching by passing a `CacheMode` value to `CrawlerRunConfig` from [`crawl4ai/async_configs.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/async_configs.py):

```python
import asyncio
from crawl4ai import AsyncWebCrawler, CacheMode
from crawl4ai.async_configs import CrawlerRunConfig

async def warm_cache():
    # Populate cache without reading stale entries

    config = CrawlerRunConfig(cache_mode=CacheMode.WRITE_ONLY)
    async with AsyncWebCrawler(verbose=True) as crawler:
        await crawler.arun(
            url="https://example.com/catalog",
            config=config
        )

async def validate_from_cache():
    # Read existing data but never update it

    config = CrawlerRunConfig(cache_mode=CacheMode.READ_ONLY)
    async with AsyncWebCrawler(verbose=True) as crawler:
        result = await crawler.arun(
            url="https://example.com/catalog",
            config=config
        )
        print(result.markdown)

asyncio.run(warm_cache())

```

## Migrating from Legacy Boolean Flags

Earlier versions of crawl4ai used fragmented boolean flags such as `bypass_cache`, `disable_cache`, `no_cache_read`, and `no_cache_write`. The current enum supersedes these flags with a one-to-one mapping:

| Legacy Flag | New CacheMode |
|-------------|---------------|
| `bypass_cache=True` | `CacheMode.BYPASS` |
| `disable_cache=True` | `CacheMode.DISABLED` |
| `no_cache_read=True` | `CacheMode.WRITE_ONLY` |
| `no_cache_write=True` | `CacheMode.READ_ONLY` |

Refer to [`docs/md_v2/core/cache-modes.md`](https://github.com/unclecode/crawl4ai/blob/main/docs/md_v2/core/cache-modes.md) in the repository for the complete migration guide and additional context on when to use each mode.

## Summary

- **`CacheMode.ENABLED`** reads and writes cache entries (default production behavior).
- **`CacheMode.DISABLED`** eliminates all cache I/O.
- **`CacheMode.READ_ONLY`** reads existing entries but never writes, protecting cache integrity.
- **`CacheMode.WRITE_ONLY`** writes new entries while ignoring historical data, perfect for cache warming.
- **`CacheMode.BYPASS`** skips the cache entirely for the current operation without affecting global settings.

## Frequently Asked Questions

### What is the default caching mode in crawl4ai?

**`CacheMode.ENABLED`** is the default when you instantiate `CrawlerRunConfig` without specifying a `cache_mode` parameter. This ensures that crawls automatically benefit from disk-based deduplication and offline replay unless you explicitly opt out.

### How do I prevent crawl4ai from writing to the cache directory?

Set `cache_mode=CacheMode.READ_ONLY` in your `CrawlerRunConfig`. This guarantees that the crawler will load existing cached pages if they exist, but will never create new files or overwrite existing ones, making it safe for read-only filesystems or validation workflows.

### Is CacheMode.BYPASS the same as CacheMode.DISABLED?

While both modes prevent reading from and writing to the cache for the current operation, **`BYPASS`** is specifically designed for temporary overrides where you need fresh data while preserving the cache for other concurrent or future operations. **`DISABLED`** is a more absolute setting typically used when you want to ensure no cache files are created or consulted for the entire session.

### Where can I find the CacheMode enum in the source code?

The enum is defined in **[`crawl4ai/cache_context.py`](https://github.com/unclecode/crawl4ai/blob/main/crawl4ai/cache_context.py)** at lines 4-21, with the logic that interprets these modes implemented in the `CacheContext` class methods `should_read()` and `should_write()` at lines 59-88.