How to Validate LLM Outputs with strict_eq Equivalence Principle in GenLayer

Use gl.eq_principle.strict_eq to enforce byte‑for‑byte equality between leader and validator execution, ensuring deterministic consensus on nondeterministic LLM or web‑fetched data.

GenLayer contracts can invoke nondeterministic operations—such as LLM prompts via gl.nondet.exec_prompt or web fetches via gl.nondet.web.render—while still achieving deterministic consensus through an equivalence principle. The strict_eq principle mandates that the leader's output must be exactly identical to the validator's output, making it essential for validating LLM outputs in production contracts.

What strict_eq Guarantees

strict_eq provides the strongest form of equivalence checking available in the GenLayer SDK:

  • Exact byte‑for‑byte equality of the JSON‑serialized result
  • Consistent key ordering (achieved via sort_keys=True in json.dumps)
  • Automatic validator generation — the SDK creates a sandboxed sub‑VM that re‑executes the leader call

When consensus fails, the transaction is rejected preventing divergent state across validator nodes.

Using strict_eq in Production Contracts

The canonical pattern wraps a leader function containing nondeterministic calls with gl.eq_principle.strict_eq:


# contracts/football_bets.py

def _check_match(self, resolution_url: str, team1: str, team2: str) -> dict:
    def get_match_result() -> str:
        web_data = gl.nondet.web.render(resolution_url, mode="text")
        task = f"""
Extract the match result for:
Team 1: {team1}
Team 2: {team2}
Web content:
{web_data}
Respond in JSON:
{{"score": str, "winner": int}}
"""
        result = gl.nondet.exec_prompt(task, response_format="json")
        return json.dumps(result, sort_keys=True)

    # strict_eq guarantees that every validator obtains the exact same JSON string

    result_json = json.loads(gl.eq_principle.strict_eq(get_match_result))
    return result_json

Critical implementation detail: Serialize with sort_keys=True to ensure deterministic key ordering. Without this, semantically identical JSON objects may fail strict_eq validation due to key order differences.

Source: [contracts/football_bets.py](https://github.com/genlayerlabs/genlayer-project-boilerplate/blob/main/contracts/football_bets.py#L29-L55)

Testing Limitations and Workarounds

Why strict_eq Fails in Direct Mode

strict_eq internally invokes spawn_sandbox to create an isolated execution environment for validators. The mock VM used in direct-mode tests does not implement sandbox calls, causing strict_eq to fail with unsupported operation errors.

Alternative: run_nondet_unsafe for Unit Tests

Use glvm.run_nondet_unsafe.lazy to manually define validator logic without sandbox dependency:


# tests/integration/test_new_features.py

class NondetContract(gl.Contract):
    last_result: str

    @gl.public.write
    def fetch_and_store(self, url: str) -> str:
        def leader() -> str:
            resp = gl.nondet.web.get(url)
            return resp.body.decode("utf-8", errors="replace")

        def validator(result: glvm.Result) -> bool:
            if not isinstance(result, glvm.Return):
                return False
            resp = gl.nondet.web.get(url)
            my_val = resp.body.decode("utf-8", errors="replace")
            return my_val == result.calldata

        value = glvm.run_nondet_unsafe.lazy(leader, validator).get()
        self.last_result = value
        return value

Source: [tests/integration/test_new_features.py](https://github.com/genlayerlabs/genlayer-project-boilerplate/blob/main/tests/integration/test_new_features.py#L48-L73)

Running Validators in Direct Mode Tests

The direct_vm fixture captures registered validators for explicit execution:

def test_run_validator_true_when_mocks_agree(direct_vm, direct_deploy, nondet_contract_path):
    direct_vm.mock_web(r".*api\.example\.com.*", {"method":"GET","status":200,"body":b"result_A"})
    contract = direct_deploy(nondet_contract_path)
    contract.fetch_and_store("https://api.example.com/data")
    assert direct_vm.run_validator() is True

Source: [tests/integration/test_new_features.py](https://github.com/genlayerlabs/genlayer-project-boilerplate/blob/main/tests/integration/test_new_features.py#L64-L81)

strict_eq vs. run_nondet_unsafe Comparison

Approach Use Case Sandbox Required Test Compatibility
gl.eq_principle.strict_eq Production contracts requiring guaranteed consensus Yes Integration tests only
glvm.run_nondet_unsafe.lazy Unit tests with mocked dependencies No Direct and integration tests

Best Practices for LLM Output Validation

  1. Always sort JSON keys when serializing results for strict_eq — use json.dumps(result, sort_keys=True)

  2. Design prompts for determinism — include explicit formatting instructions and temperature controls where supported

  3. Use integration tests for strict_eq — deploy to a full test environment with sandbox support rather than direct-mode mocks

  4. Implement fallback validation for run_nondet_unsafe — validate both return type and semantic equality, not just raw bytes

Summary

  • strict_eq enforces exact output equality between leader and validator execution for nondeterministic operations
  • spawn_sandbox enables isolated validator re‑execution but is unavailable in direct‑mode tests
  • run_nondet_unsafe.lazy provides explicit validator control for unit testing without sandbox dependencies
  • Production contracts using LLM outputs should combine strict_eq with deterministic serialization (sort_keys=True)

Frequently Asked Questions

What happens if validator output differs from leader output in strict_eq?

The transaction is rejected. According to the GenLayer protocol, consensus failure occurs when strict_eq detects any byte‑level difference between the leader's captured result and the validator's re‑executed result. No partial acceptance or fuzzy matching is applied.

Can I use strict_eq with non‑JSON LLM outputs?

Yes, though JSON is recommended for structured data. strict_eq performs raw string comparison, so any serializable format works. Ensure consistent encoding and apply any necessary normalization (whitespace trimming, case normalization) within the leader function before returning.

Why does my strict_eq test fail in pytest but pass in production?

Most likely you're running in direct mode where spawn_sandbox is unimplemented. Switch to integration test configuration or refactor to use run_nondet_unsafe with explicit validator functions. Check your conftest.py to confirm which VM mode your tests invoke.

How do I debug validator failures in strict_eq?

Add logging within your leader function to capture the exact serialized string being compared. Since strict_eq uses spawn_sandbox, you cannot step through validator execution directly—instead, log inputs and outputs at the leader boundary and compare against validator re‑execution in a separate debug environment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →