How to Mock LLM Responses in GenLayer Direct Tests: A Complete Guide
GenLayer's direct-mode VM fixture provides a mock_llm method that accepts a regex pattern and a response string, allowing you to intercept gl.nondet.exec_prompt calls and return deterministic stub data instead of invoking actual LLM services.
When testing GenLayer contracts locally, you need deterministic control over non-deterministic dependencies like Large Language Models. The genlayer-project-boilerplate repository includes a specialized direct-mode testing framework that runs contracts in-memory while letting you mock LLM responses using pattern-based interception.
Understanding the Direct-Mode VM Fixture
The direct-mode VM is a testing fixture provided in tests/direct/conftest.py that creates an isolated, in-memory execution environment for your GenLayer contracts. Unlike integration tests that call external services, this fixture intercepts all gl.nondet.exec_prompt calls at the VM level.
According to the source code in tests/direct/conftest.py, the direct_vm fixture exposes several critical methods for test control:
mock_llm(regex, response)– Registers a pattern-to-response mapping for LLM callsmock_web(regex, response)– Provides similar functionality for HTTP requestsclear_mocks()– Removes all registered mocks to prevent test cross-contaminationexpect_revert(message)– Context manager for testing failure conditions
When your contract calls gl.nondet.exec_prompt, the direct-mode VM checks the prompt against registered regex patterns and returns the matching stub response immediately.
The mock_llm Method Syntax and Parameters
The mock_llm method signature implemented in the direct-mode VM follows this structure:
direct_vm.mock_llm(regex_pattern, response_string_or_json)
Parameter breakdown:
regex_pattern– A Python regular expression string that matches against the full prompt sent togl.nondet.exec_prompt. The pattern uses standardremodule syntax and can match partial strings using.*wildcards.response_string_or_json– The exact string returned to the contract. For JSON responses, usejson.dumps()to serialize Python dictionaries. For text responses, pass the raw string directly.
Mocks persist across contract calls until explicitly cleared. This allows multiple operations to use the same stub, but requires careful cleanup between test scenarios.
Step-by-Step Implementation Examples
Basic LLM Mocking Pattern
The most common pattern, as demonstrated in tests/direct/test_views.py (lines 38-41), involves mocking a contract method that extracts structured data:
import json
def test_simple_llm_mock(direct_vm, direct_deploy, direct_alice):
contract = direct_deploy("contracts/football_bets.py")
direct_vm.sender = direct_alice
# Register a mock matching any prompt containing "Extract the match result"
direct_vm.mock_llm(
r".*Extract the match result.*",
json.dumps({"score": "2:1", "winner": 1}),
)
# This contract method internally calls gl.nondet.exec_prompt
result = contract.get_match_result(True)
assert result["score"] == "2:1"
assert result["winner"] == 1
The regex r".*Extract the match result.*" ensures the mock triggers regardless of surrounding context in the prompt, while json.dumps() ensures the contract receives valid JSON that its parser can handle.
Testing Multiple Scenarios with clear_mocks()
When testing sequential operations in the same function, you must reset mocks to avoid contamination. The clear_mocks() method, shown in tests/direct/test_views.py (lines 44-46), removes all registered interceptors:
def test_multiple_resolves(direct_vm, direct_deploy, direct_alice):
contract = direct_deploy("contracts/football_bets.py")
direct_vm.sender = direct_alice
alice_address = to_hex(direct_alice)
# First scenario: Mock a winning result
contract.create_bet("2024-06-20", "Spain", "Italy", "1")
direct_vm.mock_llm(
r".*Extract the match result.*",
json.dumps({"score": "1:0", "winner": 1})
)
contract.resolve_bet("2024-06-20_spain_italy")
assert contract.get_player_points(alice_address) == 1
# Critical: Clear mocks before changing the scenario
direct_vm.clear_mocks()
# Second scenario: Mock a draw result
contract.create_bet("2024-06-21", "Denmark", "England", "0")
direct_vm.mock_llm(
r".*Extract the match result.*",
json.dumps({"score": "1:1", "winner": 0})
)
contract.resolve_bet("2024-06-21_denmark_england")
assert contract.get_player_points(alice_address) == 2
Failure to call clear_mocks() between scenarios can cause earlier mocks to match unexpectedly, leading to flaky tests that pass or fail based on registration order.
Mocking Non-JSON Text Responses
Not all LLM interactions require structured data. For contracts expecting raw text analysis or summaries, pass the string directly without JSON encoding:
def test_text_summary(direct_vm, direct_deploy):
contract = direct_deploy("contracts/football_bets.py")
direct_vm.mock_llm(
r".*Summarize the match.*",
"Team 1 dominated possession with 65% control and scored twice in the second half."
)
summary = contract.get_match_summary("2024-06-20")
assert "dominated possession" in summary
The direct-mode VM returns the exact byte string provided, preserving whitespace and special characters exactly as written.
Handling Contract Reverts with Invalid Mocks
Test your contract's error handling by providing malformed LLM responses and asserting that the contract reverts properly. The expect_revert context manager validates that the contract raises an exception with the expected message:
def test_invalid_llm_response(direct_vm, direct_deploy, direct_alice):
contract = direct_deploy("contracts/football_bets.py")
direct_vm.sender = direct_alice
# Provide intentionally invalid JSON missing required fields
direct_vm.mock_llm(
r".*Extract the match result.*",
"{}"
)
with direct_vm.expect_revert("Invalid match data"):
contract.resolve_bet("2024-06-20_spain_italy")
This pattern ensures your contract validates LLM outputs defensively before processing them, preventing crashes from unexpected API responses.
Creating Reusable Mock Helpers
For complex test suites, extract mock setup into helper functions to keep tests DRY. The test_resolve_bet.py file demonstrates this pattern with a private helper that configures both web and LLM mocks simultaneously:
def _setup_match_mocks(vm, score, winner):
"""Configure both data source mocks for a match resolution test."""
vm.mock_web(
r".*bbc\.com/sport/football/scores-fixtures.*",
{"status": 200, "body": f"Result: {score}, winner: {winner}"},
)
vm.mock_llm(
r".*Extract the match result.*",
json.dumps({"score": score, "winner": winner}),
)
# Usage in test functions
def test_specific_match(direct_vm, direct_deploy):
contract = direct_deploy("contracts/football_bets.py")
_setup_match_mocks(direct_vm, "3:2", 1)
result = contract.resolve_bet("2024-06-20_spain_italy")
assert result.is_completed
Helper functions like this centralize your mock data, making updates easier when contract requirements change.
Summary
- The
direct_vmfixture fromtests/direct/conftest.pyprovides an in-memory VM that interceptsgl.nondet.exec_promptcalls during testing. - Use
direct_vm.mock_llm(regex, response)to register pattern-based stubs that return deterministic data instead of calling real LLMs. - Pass
json.dumps()for structured responses or raw strings for text outputs as the second parameter. - Always call
direct_vm.clear_mocks()between test scenarios to prevent pattern matching conflicts. - Test error handling by providing invalid mock data and asserting reverts with
direct_vm.expect_revert(). - Extract complex mock setups into helper functions to maintain clean, readable test files.
Frequently Asked Questions
What is the direct_vm fixture in GenLayer?
The direct_vm fixture is a specialized testing utility defined in tests/direct/conftest.py that creates an isolated, in-memory virtual machine for executing GenLayer contracts. Unlike standard test environments, this fixture intercepts non-deterministic operations like LLM calls and web requests, allowing you to register mock responses using methods like mock_llm() and control execution context through properties like sender.
How do I clear LLM mocks between test scenarios?
Call direct_vm.clear_mocks() to remove all registered mocks from the VM instance. This method is essential when testing multiple contract operations in a single test function, as documented in tests/direct/test_views.py lines 44-46. Without clearing, previously registered regex patterns might match subsequent prompts unexpectedly, causing test failures or false positives.
Can I mock both LLM and web calls in the same test?
Yes. The direct-mode VM supports simultaneous mocking of multiple dependency types. Use direct_vm.mock_llm() for gl.nondet.exec_prompt calls and direct_vm.mock_web() for HTTP requests within the same test function. Both methods accept regex patterns and responses, allowing you to simulate complete data pipelines where web scraping feeds data to LLM analysis, as shown in the _setup_match_mocks helper pattern from tests/direct/test_resolve_bet.py.
What format should LLM mock responses use?
The format depends entirely on your contract's expectations. If the contract parses JSON using json.loads() or similar, pass json.dumps({"key": "value"}) as the response parameter. If the contract expects raw text, pass the string directly. The direct-mode VM performs no transformation on the response—it returns the exact bytes provided, so matching your contract's parsing logic is critical for successful tests.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →