How to Wait for JavaScript to Execute and Handle Dynamic Content with crawl4ai
Use the wait_for parameter in CrawlerRunConfig to pause execution until a CSS selector appears or a JavaScript expression returns true, ensuring dynamic content is fully rendered before extraction.
The crawl4ai library renders web pages using a real Playwright-controlled browser, making it ideal for scraping modern JavaScript-heavy applications. When dealing with dynamic content that loads asynchronously, you need a reliable mechanism to pause crawling until specific conditions are met. The library provides this through the wait_for configuration option, implemented primarily in crawl4ai/async_crawler_strategy.py.
Understanding the wait_for Mechanism
The core waiting logic resides in the smart_wait method within crawl4ai/async_crawler_strategy.py (lines 231-290). This method intelligently interprets your wait condition and delegates to the appropriate Playwright API or custom JavaScript evaluator.
How smart_wait Detects Wait Types
The method analyzes the wait_for string to determine the waiting strategy:
js:prefix — Treats the remainder as a JavaScript snippet to evaluate repeatedly until it returnstruecss:prefix — Treats the remainder as a CSS selector to wait for viapage.wait_for_selector()- No prefix — First attempts to parse as a JavaScript function (detecting
() =>orfunctionsyntax), then falls back to CSS selector if parsing fails
CSP-Compliant JavaScript Execution
For JavaScript waits, smart_wait calls csp_compliant_wait (lines 294-314 in crawl4ai/async_crawler_strategy.py). This helper repeatedly evaluates your JavaScript expression every 100 milliseconds until it returns a truthy value or the timeout expires (defaulting to the page_timeout value, approximately 60 seconds).
If the condition never becomes true, the method raises a clear ValueError indicating the wait condition timed out, making debugging straightforward.
Waiting for CSS Selectors
Use the css: prefix when you need to wait for a specific element to appear in the DOM. This leverages Playwright's built-in wait_for_selector method, which is efficient and handles mutation observation automatically.
import asyncio
from crawl4ai import AsyncWebCrawler, BrowserConfig
async def main():
browser_cfg = BrowserConfig(headless=True)
async with AsyncWebCrawler(config=browser_cfg) as crawler:
# Wait until an element with class "article" appears
result = await crawler.arun(
url="https://example.com/news",
wait_for="css:.article",
screenshot=True
)
print(result.markdown)
asyncio.run(main())
In this example, smart_wait detects the css: prefix and calls page.wait_for_selector('.article'), pausing execution until the element is present or the timeout occurs.
Waiting for JavaScript Conditions
For complex dynamic content that doesn't correspond to simple DOM elements—such as waiting for specific data to load, animations to complete, or custom JavaScript variables to be set—use the js: prefix.
import asyncio
from crawl4ai import AsyncWebCrawler
async def main():
async with AsyncWebCrawler() as crawler:
# Wait until at least 5 product elements are loaded
result = await crawler.arun(
url="https://example.com/shop",
wait_for="js:return document.querySelectorAll('.product').length >= 5"
)
print(result.markdown)
asyncio.run(main())
Here, csp_compliant_wait repeatedly evaluates the JavaScript expression every 100ms until it returns true (indicating five or more products are present) or the default timeout expires.
Implicit Detection Without Prefixes
You can omit prefixes entirely, and smart_wait will attempt to infer the correct type. This is useful for quick scripts but requires careful syntax to avoid ambiguity.
import asyncio
from crawl4ai import AsyncWebCrawler
async def main():
async with AsyncWebCrawler() as crawler:
# No prefix: treated as JS due to arrow function syntax
result = await crawler.arun(
url="https://example.com/gallery",
wait_for="() => window.galleryReady === true"
)
print(result.markdown)
asyncio.run(main())
The method detects the () => pattern and routes this to csp_compliant_wait. If the string fails JavaScript parsing, it falls back to treating the string as a CSS selector.
Using the Synchronous API
The wait_for parameter works identically in the synchronous WebCrawler class, which wraps the same underlying logic for simpler scripting scenarios.
from crawl4ai import WebCrawler, BrowserConfig
crawler = WebCrawler(config=BrowserConfig())
result = crawler.run(
url="https://example.com",
wait_for="css:#main-content"
)
print(result.markdown)
Both async and sync implementations delegate to the same smart_wait logic in crawl4ai/async_crawler_strategy.py, ensuring consistent behavior across APIs.
Summary
- crawl4ai handles dynamic content through the
wait_forparameter inCrawlerRunConfig, implemented incrawl4ai/async_crawler_strategy.py. - Use
css:prefix to wait for DOM elements via Playwright'swait_for_selector. - Use
js:prefix to wait for custom JavaScript conditions evaluated bycsp_compliant_waitin 100ms polling intervals. - Omit prefixes to let
smart_waitauto-detect JavaScript functions (detecting() =>orfunctionsyntax) or fall back to CSS selectors. - Both async (
AsyncWebCrawler) and sync (WebCrawler) APIs support identicalwait_forfunctionality with default timeouts around 60 seconds.
Frequently Asked Questions
What is the default timeout for wait_for in crawl4ai?
The default timeout follows the page_timeout configuration value, which is approximately 60 seconds. If the condition (CSS selector or JavaScript expression) is not satisfied within this window, smart_wait raises a ValueError indicating the wait condition timed out. You can adjust this timeout by modifying the page_timeout parameter in your BrowserConfig.
Can I wait for multiple conditions simultaneously?
Currently, wait_for accepts a single string condition. To wait for multiple conditions, you should combine them into a single JavaScript expression using logical operators. For example: wait_for="js:return document.querySelector('.products') && document.querySelector('.reviews')". This ensures both elements exist before proceeding with content extraction.
How does crawl4ai handle Content Security Policy (CSP) restrictions when executing JavaScript waits?
The library uses csp_compliant_wait (lines 294-314 in crawl4ai/async_crawler_strategy.py), which evaluates JavaScript within the page context using Playwright's page.evaluate(). This approach respects the page's CSP because it runs inside the browser's normal execution environment rather than injecting external scripts. The method polls every 100 milliseconds until the condition returns true or the timeout expires.
Is there a performance difference between CSS and JavaScript waits?
Yes. CSS selector waits leverage Playwright's native wait_for_selector, which uses browser-level mutation observers and is generally more efficient with lower overhead. JavaScript waits require repeated execution of page.evaluate every 100ms via csp_compliant_wait, consuming slightly more CPU and network resources. For simple element presence, prefer css: selectors; use js: only when you need complex logic like checking element text content, computed styles, or global state variables.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →