# How to Evaluate the Output of DeepSeek-Reasonix: A Contract-Driven Validation System

> Learn how to evaluate DeepSeek-Reasonix output using its contract-driven architecture. Discover how tools report status and provide verifiable evidence for Guardian review.

- Repository: [YHH/DeepSeek-Reasonix](https://github.com/esengine/DeepSeek-Reasonix)
- Tags: how-to-guide
- Published: 2026-08-10

---

**DeepSeek-Reasonix evaluates every output through a layered contract-driven architecture that requires tools to report status via `update_goal` and provide verifiable evidence through `complete_step`, validated by the Guardian reviewer and permission engine before acceptance.**

The open-source repository `esengine/DeepSeek-Reasonix` implements a rigorous framework to evaluate DeepSeek-Reasonix output through mandatory tool contracts and evidence-based signing. Unlike traditional LLM systems that generate responses without structured verification, Reasonix forces every turn to conclude with explicit status reporting and evidence submission, ensuring that every claim is backed by verifiable data and safety-checked before acceptance.

## Understanding the Contract-Driven Evaluation Architecture

Reasonix employs a five-layer validation stack defined in the source code. Each layer inspects output before it reaches the user or session log.

### Tool Contract Layer

At the foundation, every built-in tool registers a JSON schema in [`internal/tool/contract.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/tool/contract.go). The **tool contract** declares parameters, a read-only flag, and descriptions. The host validates every tool call against this schema using `tool.BuiltinContractEntries` before execution proceeds.

### Output-Verification Tools

Two specialized tools handle evaluation reporting:

- **`update_goal`** (implemented in [`internal/tool/builtin/updategoal.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/tool/builtin/updategoal.go)): Reports whether the current goal status is `continue`, `complete`, or `blocked`, optionally including a reason string.
- **`complete_step`** (implemented in [`internal/tool/builtin/completestep.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/tool/builtin/completestep.go)): Records a signed-off step with mandatory evidence. The host rejects any claim lacking at least one evidence item (such as a diff, verification output, or manual check).

### Guardian Safety Reviewer

Before any tool call executes, the **Guardian** reviewer inspects the payload for prohibited content. Defined in [`internal/guardian/guardian.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/guardian/guardian.go), this LLM-driven safety layer can reject unsafe outputs even if they pass schema validation.

### Permission Engine

The **permission engine** in [`internal/permission/permission.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/permission/permission.go) evaluates whether a call complies with policy rules. The `permission.Decide` function checks `ask`, `require`, and `prefer` rules against the current goal context before allowing execution.

### Session and Memory Integration

Evidence from `complete_step` is integrated into the session log via [`internal/agent/run_loop.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/agent/run_loop.go), enabling replay and rewind functionality while maintaining an immutable record of verified outputs.

## The Four-Step Evaluation Process

When Reasonix produces output, the host executes this validation sequence:

1. **Validate the tool contract** — The JSON payload must match the schema generated by `tool.BuiltinContractEntries`.
2. **Run the Guardian** — The safety reviewer checks for prohibited content in the pending tool call.
3. **Apply permission rules** — The system evaluates whether the call complies with explicit policy constraints.
4. **Record the evidence** — Successful calls append to the session log with full evidence chains.

If any step fails, the host returns an error and rejects the turn, forcing the model to retry or report blockage.

## Practical Implementation Examples

### Evaluating Output via CLI

When using the interactive REPL, you explicitly close evaluation loops:

```bash

# Start a Reasonix session

reasonix

# Execute a command that produces output

> bash "git status"

# Report turn disposition

> update_goal {"status":"continue","reason":"need to add files","next_action":"git add ."}

# Sign off the step with evidence

> complete_step {"step":"Check repository status","result":"untracked files found","evidence":[{"kind":"verification","summary":"git status output"}]}

```

### Programmatic Evaluation with the Go SDK

The Go SDK enforces the same validation layers when evaluating output programmatically:

```go
import (
    "context"
    "github.com/esengine/DeepSeek-Reasonix/sdk/go"
)

func main() {
    client := reasonix.NewClient()

    // Execute tool and capture output
    out, _ := client.CallTool(context.Background(), reasonix.Tool{
        Name: "bash",
        Args: map[string]any{
            "command": "go test ./...",
        },
    })

    // Evaluate: Report completion status
    client.CallTool(context.Background(), reasonix.Tool{
        Name: "update_goal",
        Args: map[string]any{
            "status": "complete",
            "reason": "All tests passed",
        },
    })

    // Sign off with evidence
    client.CallTool(context.Background(), reasonix.Tool{
        Name: "complete_step",
        Args: map[string]any{
            "step":   "Run unit tests",
            "result": "passed",
            "evidence": []any{
                map[string]any{
                    "kind":    "verification",
                    "summary": out,
                },
            },
        },
    })
}

```

### Custom Plugin Evaluation

Custom MCP plugins must also report evidence through the evaluation pipeline:

```go
func MyPluginAction(ctx context.Context, args map[string]any) (any, error) {
    result := "generated code"

    // Register evidence with Reasonix
    _, err := reasonix.CallTool(ctx, reasonix.Tool{
        Name: "complete_step",
        Args: map[string]any{
            "step":   "Generate code",
            "result": "ok",
            "evidence": []any{
                map[string]any{
                    "kind":    "manual",
                    "summary": "Human verified the generated code",
                },
            },
        },
    })
    return result, err
}

```

The Guardian and permission engine evaluate these plugin calls identically to built-in tools.

## Summary

- **Contract validation** in [`internal/tool/contract.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/tool/contract.go) ensures all tool calls match declared JSON schemas.
- **Status reporting** via `update_goal` and evidence submission via `complete_step` are mandatory for finalizing outputs.
- **Safety review** by the Guardian in [`internal/guardian/guardian.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/guardian/guardian.go) filters prohibited content before acceptance.
- **Permission checks** through `permission.Decide` enforce policy compliance based on goal context.
- **Session logging** in [`internal/agent/run_loop.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/agent/run_loop.go) maintains immutable evidence chains for audit and replay.

## Frequently Asked Questions

### What happens if I don't call `complete_step` after a tool execution?

The host will reject the turn and return an error. Reasonix requires either an `update_goal` report for every turn, and `complete_step` with at least one evidence item when finishing planned steps. Without this verification, the output remains in a pending state and is not committed to the session log.

### Can the Guardian reviewer reject valid tool calls?

Yes. The Guardian in [`internal/guardian/guardian.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/guardian/guardian.go) operates independently of schema validation. Even if a tool call matches its JSON contract and passes permission checks, the safety reviewer can reject it for containing unsafe content or violating security policies, forcing the model to revise its approach.

### What types of evidence does `complete_step` accept?

According to [`internal/tool/builtin/completestep.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/tool/builtin/completestep.go), evidence items must include a `kind` field (such as `diff`, `verification`, or `manual`) and a `summary` string. The host rejects any `complete_step` call that lacks at least one evidence item, ensuring all reported results are verifiable.

### How does the permission engine handle conflicting rules?

The `permission.Decide` function in [`internal/permission/permission.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/permission/permission.go) evaluates rules in the context of the current goal and any explicit "ask" configurations. It applies `require`, `prefer`, and `ask` directives hierarchically, denying execution if mandatory constraints are violated while allowing optional preferences to guide selection when multiple valid paths exist.