How to Render Web Pages for Data Extraction in GenLayer Contracts

GenLayer contracts fetch and process live web content using the gl.nondet.web.render API, which retrieves pages as text or screenshots for LLM processing, and must be wrapped in equivalence-principle validators to ensure validator consensus.

The genlayerlabs/genlayer-project-boilerplate repository provides production-ready patterns for accessing real-time web data directly from smart contracts. By combining the non-deterministic web rendering capabilities with LLM-based parsing, developers can extract structured information from external websites while maintaining the security guarantees of the GenLayer protocol.

Using the Non-Deterministic Web Render API

The core entry point for fetching web content is the gl.nondet.web.render function. According to the GenLayer SDK source, this function accepts a URL and a mode parameter to determine the output format.

Function Signature and Modes

The complete function signature is:

gl.nondet.web.render(url: str, *, mode: Literal['text', 'html', 'screenshot']) → Response

The mode parameter supports three distinct output formats:

  • text: Returns the raw HTML content as a string, suitable for parsing text-based data
  • html: Returns the page markup (alternative to text mode)
  • screenshot: Returns an Image object containing a PNG screenshot of the rendered page

Text Mode for Structured Data Extraction

When operating in text mode, the function retrieves the raw HTML of the target page. This mode is ideal for extracting structured data like sports scores, prices, or article content.

In contracts/football_bets.py, the contract renders a match results page and feeds the HTML to an LLM for extraction:

def _check_match(self, resolution_url: str, team1: str, team2: str) -> dict:
    # Render the page as plain text (raw HTML)

    web_data = gl.nondet.web.render(resolution_url, mode="text")
    
    # Build extraction prompt

    task = f"""
    Extract the match result for:
    Team 1: {team1}
    Team 2: {team2}
    
    Web content:
    {web_data}
    
    Respond in JSON:
    {{
        "score": str,
        "winner": int
    }}
    """
    
    # Execute LLM prompt with enforced JSON output

    result = gl.nondet.exec_prompt(task, response_format="json")
    return json.loads(result)

Screenshot Mode for Visual Analysis

For pages where visual layout matters or text extraction is insufficient, screenshot mode captures the rendered page as an image:

@gl.public.write
def describe_screenshot(self, url: str) -> str:
    # Render the page as a PNG screenshot

    img = gl.nondet.web.render(url, mode="screenshot")
    
    # Pass image bytes to LLM for analysis

    description = gl.nondet.exec_prompt(
        "Describe this image", 
        images=[img.raw]
    )
    
    self.last_desc = description
    return self.last_desc

Screenshot mode returns an Image object with a raw property containing the PNG bytes, enabling OCR, visual verification, or image description workflows.

Integrating with the Equivalence Principle

Because web calls are inherently non-deterministic, GenLayer contracts must wrap all gl.nondet.web.render calls within equivalence-principle validators to guarantee consensus across validator nodes.

Using strict_eq for Deterministic Consensus

The standard pattern recorded in the boilerplate uses gl.eq_principle.strict_eq to ensure all validators execute the same web call and reach identical conclusions. As implemented in contracts/football_bets.py at line 31, the validated pattern looks like this:

result = gl.nondet.exec_prompt(task, response_format="json")

# Validate with strict_eq so validators see the same result

result_json = json.loads(gl.eq_principle.strict_eq(lambda: result))

The strict_eq validator records the lambda function and re-executes it on each validator node, comparing outputs to ensure they match exactly before the transaction is finalized.

Processing Returned Content with LLMs

After rendering a page, contracts typically use gl.nondet.exec_prompt to process the content. The workflow follows three steps:

  1. Render: Fetch the page content using gl.nondet.web.render
  2. Prompt: Construct a task that includes the web data and specifies the desired output format
  3. Parse: Force JSON output using response_format="json" and parse with json.loads

This pattern ensures structured data extraction from unstructured HTML, as demonstrated in the football bets contract where match scores and winner IDs are extracted from sports websites.

Testing with Mocked Web Calls

The GenLayer SDK provides deterministic testing through the direct_vm fixture, allowing developers to intercept web calls without making actual HTTP requests.

Mocking Text Responses

In direct-mode tests, use direct_vm.mock_web to intercept HTTP requests based on URL patterns:

def test_match_extraction(direct_vm, direct_deploy):
    # Mock the HTTP response for BBC sports pages

    direct_vm.mock_web(r".*bbc.*", {
        "status": 200, 
        "body": "Spain 3-0 Italy"
    })
    
    # Mock the LLM response

    direct_vm.mock_llm(r".*", '{"score": "3:0", "winner": 1}')
    
    contract = direct_deploy("contracts/football_bets.py")
    contract.create_bet("2024-06-20", "Spain", "Italy", "1")
    contract.resolve_bet("2024-06-20_spain_italy")
    
    # Verify extracted data

    bet = contract.get_bets()[0]
    assert bet.real_score == "3:0"

The mock_web helper is defined in tests/direct/conftest.py and supports both text and screenshot interception patterns.

Handling Screenshot Mock Limitations

When testing screenshot mode, the integration test suite in tests/integration/test_new_features.py reveals a current limitation. At line 97, the test demonstrates screenshot interception, but as documented at line 61, the underlying wasi_mock returns empty bytes when mocking screenshots. This causes a PIL.UnidentifiedImageError when the contract attempts to process the mocked image data.

Developers should account for this limitation when writing tests for image-based extraction workflows.

Security and Determinism Considerations

All non-deterministic operations in GenLayer must adhere to the equivalence principle to maintain blockchain consensus. When rendering web pages:

  • Always wrap gl.nondet.web.render calls inside gl.eq_principle.strict_eq or other validators
  • Canonicalize results before persistence by using sorted JSON keys or normalized string formats
  • Validate LLM outputs to ensure consistent parsing across different validator executions

Failure to wrap web calls properly results in non-deterministic state that prevents validators from reaching consensus, causing transaction rejection.

Summary

  • Use gl.nondet.web.render with mode="text" for HTML extraction or mode="screenshot" for visual analysis
  • Wrap all web calls in gl.eq_principle.strict_eq validators to ensure cross-validator consensus
  • Process content using gl.nondet.exec_prompt with response_format="json" for structured data extraction
  • Test deterministically using direct_vm.mock_web to intercept HTTP calls in direct-mode tests
  • Account for limitations when mocking screenshots due to empty byte returns in the current wasi_mock implementation
  • Refer to contracts/football_bets.py for a complete production example of text-based web extraction

Frequently Asked Questions

What is the difference between text and screenshot modes in GenLayer contracts?

Text mode retrieves the raw HTML content of a page as a string, which is optimal for extracting textual data like prices, scores, or article content. Screenshot mode captures the visually rendered page as a PNG image object, enabling visual analysis, OCR, or layout-dependent extraction. Both modes require wrapping in strict_eq validators to ensure deterministic consensus.

Why must web render calls be wrapped in equivalence-principle validators?

Web calls are non-deterministic because external websites can change between executions or return different content to different validators. The gl.eq_principle.strict_eq validator ensures all validator nodes execute the same web call and produce identical results before the smart contract state is updated, maintaining the deterministic requirements of blockchain consensus.

How do you test contracts that use web rendering without making real HTTP calls?

Use the direct_vm fixture from tests/direct/conftest.py and call direct_vm.mock_web(pattern, response) to intercept HTTP requests. The pattern is a regex matching the target URL, and the response dictionary contains the status code and body text. For LLM calls, use direct_vm.mock_llm(pattern, response) to return deterministic JSON outputs during testing.

What are the limitations when mocking screenshot renders in tests?

When using direct_vm.mock_web to intercept screenshot mode calls, the underlying wasi_mock currently returns empty bytes for the image data, causing a PIL.UnidentifiedImageError when the contract attempts to open the image. This limitation is documented in tests/integration/test_new_features.py at lines 61 and 97, and developers should handle this exception or test screenshot logic separately from the mock interception.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →