How to Render Web Pages for Data Extraction in GenLayer Contracts
GenLayer contracts fetch and process live web content using the gl.nondet.web.render API, which retrieves pages as text or screenshots for LLM processing, and must be wrapped in equivalence-principle validators to ensure validator consensus.
The genlayerlabs/genlayer-project-boilerplate repository provides production-ready patterns for accessing real-time web data directly from smart contracts. By combining the non-deterministic web rendering capabilities with LLM-based parsing, developers can extract structured information from external websites while maintaining the security guarantees of the GenLayer protocol.
Using the Non-Deterministic Web Render API
The core entry point for fetching web content is the gl.nondet.web.render function. According to the GenLayer SDK source, this function accepts a URL and a mode parameter to determine the output format.
Function Signature and Modes
The complete function signature is:
gl.nondet.web.render(url: str, *, mode: Literal['text', 'html', 'screenshot']) → Response
The mode parameter supports three distinct output formats:
- text: Returns the raw HTML content as a string, suitable for parsing text-based data
- html: Returns the page markup (alternative to text mode)
- screenshot: Returns an
Imageobject containing a PNG screenshot of the rendered page
Text Mode for Structured Data Extraction
When operating in text mode, the function retrieves the raw HTML of the target page. This mode is ideal for extracting structured data like sports scores, prices, or article content.
In contracts/football_bets.py, the contract renders a match results page and feeds the HTML to an LLM for extraction:
def _check_match(self, resolution_url: str, team1: str, team2: str) -> dict:
# Render the page as plain text (raw HTML)
web_data = gl.nondet.web.render(resolution_url, mode="text")
# Build extraction prompt
task = f"""
Extract the match result for:
Team 1: {team1}
Team 2: {team2}
Web content:
{web_data}
Respond in JSON:
{{
"score": str,
"winner": int
}}
"""
# Execute LLM prompt with enforced JSON output
result = gl.nondet.exec_prompt(task, response_format="json")
return json.loads(result)
Screenshot Mode for Visual Analysis
For pages where visual layout matters or text extraction is insufficient, screenshot mode captures the rendered page as an image:
@gl.public.write
def describe_screenshot(self, url: str) -> str:
# Render the page as a PNG screenshot
img = gl.nondet.web.render(url, mode="screenshot")
# Pass image bytes to LLM for analysis
description = gl.nondet.exec_prompt(
"Describe this image",
images=[img.raw]
)
self.last_desc = description
return self.last_desc
Screenshot mode returns an Image object with a raw property containing the PNG bytes, enabling OCR, visual verification, or image description workflows.
Integrating with the Equivalence Principle
Because web calls are inherently non-deterministic, GenLayer contracts must wrap all gl.nondet.web.render calls within equivalence-principle validators to guarantee consensus across validator nodes.
Using strict_eq for Deterministic Consensus
The standard pattern recorded in the boilerplate uses gl.eq_principle.strict_eq to ensure all validators execute the same web call and reach identical conclusions. As implemented in contracts/football_bets.py at line 31, the validated pattern looks like this:
result = gl.nondet.exec_prompt(task, response_format="json")
# Validate with strict_eq so validators see the same result
result_json = json.loads(gl.eq_principle.strict_eq(lambda: result))
The strict_eq validator records the lambda function and re-executes it on each validator node, comparing outputs to ensure they match exactly before the transaction is finalized.
Processing Returned Content with LLMs
After rendering a page, contracts typically use gl.nondet.exec_prompt to process the content. The workflow follows three steps:
- Render: Fetch the page content using
gl.nondet.web.render - Prompt: Construct a task that includes the web data and specifies the desired output format
- Parse: Force JSON output using
response_format="json"and parse withjson.loads
This pattern ensures structured data extraction from unstructured HTML, as demonstrated in the football bets contract where match scores and winner IDs are extracted from sports websites.
Testing with Mocked Web Calls
The GenLayer SDK provides deterministic testing through the direct_vm fixture, allowing developers to intercept web calls without making actual HTTP requests.
Mocking Text Responses
In direct-mode tests, use direct_vm.mock_web to intercept HTTP requests based on URL patterns:
def test_match_extraction(direct_vm, direct_deploy):
# Mock the HTTP response for BBC sports pages
direct_vm.mock_web(r".*bbc.*", {
"status": 200,
"body": "Spain 3-0 Italy"
})
# Mock the LLM response
direct_vm.mock_llm(r".*", '{"score": "3:0", "winner": 1}')
contract = direct_deploy("contracts/football_bets.py")
contract.create_bet("2024-06-20", "Spain", "Italy", "1")
contract.resolve_bet("2024-06-20_spain_italy")
# Verify extracted data
bet = contract.get_bets()[0]
assert bet.real_score == "3:0"
The mock_web helper is defined in tests/direct/conftest.py and supports both text and screenshot interception patterns.
Handling Screenshot Mock Limitations
When testing screenshot mode, the integration test suite in tests/integration/test_new_features.py reveals a current limitation. At line 97, the test demonstrates screenshot interception, but as documented at line 61, the underlying wasi_mock returns empty bytes when mocking screenshots. This causes a PIL.UnidentifiedImageError when the contract attempts to process the mocked image data.
Developers should account for this limitation when writing tests for image-based extraction workflows.
Security and Determinism Considerations
All non-deterministic operations in GenLayer must adhere to the equivalence principle to maintain blockchain consensus. When rendering web pages:
- Always wrap
gl.nondet.web.rendercalls insidegl.eq_principle.strict_eqor other validators - Canonicalize results before persistence by using sorted JSON keys or normalized string formats
- Validate LLM outputs to ensure consistent parsing across different validator executions
Failure to wrap web calls properly results in non-deterministic state that prevents validators from reaching consensus, causing transaction rejection.
Summary
- Use
gl.nondet.web.renderwithmode="text"for HTML extraction ormode="screenshot"for visual analysis - Wrap all web calls in
gl.eq_principle.strict_eqvalidators to ensure cross-validator consensus - Process content using
gl.nondet.exec_promptwithresponse_format="json"for structured data extraction - Test deterministically using
direct_vm.mock_webto intercept HTTP calls in direct-mode tests - Account for limitations when mocking screenshots due to empty byte returns in the current
wasi_mockimplementation - Refer to
contracts/football_bets.pyfor a complete production example of text-based web extraction
Frequently Asked Questions
What is the difference between text and screenshot modes in GenLayer contracts?
Text mode retrieves the raw HTML content of a page as a string, which is optimal for extracting textual data like prices, scores, or article content. Screenshot mode captures the visually rendered page as a PNG image object, enabling visual analysis, OCR, or layout-dependent extraction. Both modes require wrapping in strict_eq validators to ensure deterministic consensus.
Why must web render calls be wrapped in equivalence-principle validators?
Web calls are non-deterministic because external websites can change between executions or return different content to different validators. The gl.eq_principle.strict_eq validator ensures all validator nodes execute the same web call and produce identical results before the smart contract state is updated, maintaining the deterministic requirements of blockchain consensus.
How do you test contracts that use web rendering without making real HTTP calls?
Use the direct_vm fixture from tests/direct/conftest.py and call direct_vm.mock_web(pattern, response) to intercept HTTP requests. The pattern is a regex matching the target URL, and the response dictionary contains the status code and body text. For LLM calls, use direct_vm.mock_llm(pattern, response) to return deterministic JSON outputs during testing.
What are the limitations when mocking screenshot renders in tests?
When using direct_vm.mock_web to intercept screenshot mode calls, the underlying wasi_mock currently returns empty bytes for the image data, causing a PIL.UnidentifiedImageError when the contract attempts to open the image. This limitation is documented in tests/integration/test_new_features.py at lines 61 and 97, and developers should handle this exception or test screenshot logic separately from the mock interception.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →