How to Evaluate the Output of DeepSeek-Reasonix: A Contract-Driven Validation System

DeepSeek-Reasonix evaluates every output through a layered contract-driven architecture that requires tools to report status via update_goal and provide verifiable evidence through complete_step, validated by the Guardian reviewer and permission engine before acceptance.

The open-source repository esengine/DeepSeek-Reasonix implements a rigorous framework to evaluate DeepSeek-Reasonix output through mandatory tool contracts and evidence-based signing. Unlike traditional LLM systems that generate responses without structured verification, Reasonix forces every turn to conclude with explicit status reporting and evidence submission, ensuring that every claim is backed by verifiable data and safety-checked before acceptance.

Understanding the Contract-Driven Evaluation Architecture

Reasonix employs a five-layer validation stack defined in the source code. Each layer inspects output before it reaches the user or session log.

Tool Contract Layer

At the foundation, every built-in tool registers a JSON schema in internal/tool/contract.go. The tool contract declares parameters, a read-only flag, and descriptions. The host validates every tool call against this schema using tool.BuiltinContractEntries before execution proceeds.

Output-Verification Tools

Two specialized tools handle evaluation reporting:

  • update_goal (implemented in internal/tool/builtin/updategoal.go): Reports whether the current goal status is continue, complete, or blocked, optionally including a reason string.
  • complete_step (implemented in internal/tool/builtin/completestep.go): Records a signed-off step with mandatory evidence. The host rejects any claim lacking at least one evidence item (such as a diff, verification output, or manual check).

Guardian Safety Reviewer

Before any tool call executes, the Guardian reviewer inspects the payload for prohibited content. Defined in internal/guardian/guardian.go, this LLM-driven safety layer can reject unsafe outputs even if they pass schema validation.

Permission Engine

The permission engine in internal/permission/permission.go evaluates whether a call complies with policy rules. The permission.Decide function checks ask, require, and prefer rules against the current goal context before allowing execution.

Session and Memory Integration

Evidence from complete_step is integrated into the session log via internal/agent/run_loop.go, enabling replay and rewind functionality while maintaining an immutable record of verified outputs.

The Four-Step Evaluation Process

When Reasonix produces output, the host executes this validation sequence:

  1. Validate the tool contract — The JSON payload must match the schema generated by tool.BuiltinContractEntries.
  2. Run the Guardian — The safety reviewer checks for prohibited content in the pending tool call.
  3. Apply permission rules — The system evaluates whether the call complies with explicit policy constraints.
  4. Record the evidence — Successful calls append to the session log with full evidence chains.

If any step fails, the host returns an error and rejects the turn, forcing the model to retry or report blockage.

Practical Implementation Examples

Evaluating Output via CLI

When using the interactive REPL, you explicitly close evaluation loops:


# Start a Reasonix session

reasonix

# Execute a command that produces output

> bash "git status"

# Report turn disposition

> update_goal {"status":"continue","reason":"need to add files","next_action":"git add ."}

# Sign off the step with evidence

> complete_step {"step":"Check repository status","result":"untracked files found","evidence":[{"kind":"verification","summary":"git status output"}]}

Programmatic Evaluation with the Go SDK

The Go SDK enforces the same validation layers when evaluating output programmatically:

import (
    "context"
    "github.com/esengine/DeepSeek-Reasonix/sdk/go"
)

func main() {
    client := reasonix.NewClient()

    // Execute tool and capture output
    out, _ := client.CallTool(context.Background(), reasonix.Tool{
        Name: "bash",
        Args: map[string]any{
            "command": "go test ./...",
        },
    })

    // Evaluate: Report completion status
    client.CallTool(context.Background(), reasonix.Tool{
        Name: "update_goal",
        Args: map[string]any{
            "status": "complete",
            "reason": "All tests passed",
        },
    })

    // Sign off with evidence
    client.CallTool(context.Background(), reasonix.Tool{
        Name: "complete_step",
        Args: map[string]any{
            "step":   "Run unit tests",
            "result": "passed",
            "evidence": []any{
                map[string]any{
                    "kind":    "verification",
                    "summary": out,
                },
            },
        },
    })
}

Custom Plugin Evaluation

Custom MCP plugins must also report evidence through the evaluation pipeline:

func MyPluginAction(ctx context.Context, args map[string]any) (any, error) {
    result := "generated code"

    // Register evidence with Reasonix
    _, err := reasonix.CallTool(ctx, reasonix.Tool{
        Name: "complete_step",
        Args: map[string]any{
            "step":   "Generate code",
            "result": "ok",
            "evidence": []any{
                map[string]any{
                    "kind":    "manual",
                    "summary": "Human verified the generated code",
                },
            },
        },
    })
    return result, err
}

The Guardian and permission engine evaluate these plugin calls identically to built-in tools.

Summary

  • Contract validation in internal/tool/contract.go ensures all tool calls match declared JSON schemas.
  • Status reporting via update_goal and evidence submission via complete_step are mandatory for finalizing outputs.
  • Safety review by the Guardian in internal/guardian/guardian.go filters prohibited content before acceptance.
  • Permission checks through permission.Decide enforce policy compliance based on goal context.
  • Session logging in internal/agent/run_loop.go maintains immutable evidence chains for audit and replay.

Frequently Asked Questions

What happens if I don't call complete_step after a tool execution?

The host will reject the turn and return an error. Reasonix requires either an update_goal report for every turn, and complete_step with at least one evidence item when finishing planned steps. Without this verification, the output remains in a pending state and is not committed to the session log.

Can the Guardian reviewer reject valid tool calls?

Yes. The Guardian in internal/guardian/guardian.go operates independently of schema validation. Even if a tool call matches its JSON contract and passes permission checks, the safety reviewer can reject it for containing unsafe content or violating security policies, forcing the model to revise its approach.

What types of evidence does complete_step accept?

According to internal/tool/builtin/completestep.go, evidence items must include a kind field (such as diff, verification, or manual) and a summary string. The host rejects any complete_step call that lacks at least one evidence item, ensuring all reported results are verifiable.

How does the permission engine handle conflicting rules?

The permission.Decide function in internal/permission/permission.go evaluates rules in the context of the current goal and any explicit "ask" configurations. It applies require, prefer, and ask directives hierarchically, denying execution if mandatory constraints are violated while allowing optional preferences to guide selection when multiple valid paths exist.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →