How to Evaluate the Output of DeepSeek-Reasonix: A Contract-Driven Validation System
DeepSeek-Reasonix evaluates every output through a layered contract-driven architecture that requires tools to report status via update_goal and provide verifiable evidence through complete_step, validated by the Guardian reviewer and permission engine before acceptance.
The open-source repository esengine/DeepSeek-Reasonix implements a rigorous framework to evaluate DeepSeek-Reasonix output through mandatory tool contracts and evidence-based signing. Unlike traditional LLM systems that generate responses without structured verification, Reasonix forces every turn to conclude with explicit status reporting and evidence submission, ensuring that every claim is backed by verifiable data and safety-checked before acceptance.
Understanding the Contract-Driven Evaluation Architecture
Reasonix employs a five-layer validation stack defined in the source code. Each layer inspects output before it reaches the user or session log.
Tool Contract Layer
At the foundation, every built-in tool registers a JSON schema in internal/tool/contract.go. The tool contract declares parameters, a read-only flag, and descriptions. The host validates every tool call against this schema using tool.BuiltinContractEntries before execution proceeds.
Output-Verification Tools
Two specialized tools handle evaluation reporting:
update_goal(implemented ininternal/tool/builtin/updategoal.go): Reports whether the current goal status iscontinue,complete, orblocked, optionally including a reason string.complete_step(implemented ininternal/tool/builtin/completestep.go): Records a signed-off step with mandatory evidence. The host rejects any claim lacking at least one evidence item (such as a diff, verification output, or manual check).
Guardian Safety Reviewer
Before any tool call executes, the Guardian reviewer inspects the payload for prohibited content. Defined in internal/guardian/guardian.go, this LLM-driven safety layer can reject unsafe outputs even if they pass schema validation.
Permission Engine
The permission engine in internal/permission/permission.go evaluates whether a call complies with policy rules. The permission.Decide function checks ask, require, and prefer rules against the current goal context before allowing execution.
Session and Memory Integration
Evidence from complete_step is integrated into the session log via internal/agent/run_loop.go, enabling replay and rewind functionality while maintaining an immutable record of verified outputs.
The Four-Step Evaluation Process
When Reasonix produces output, the host executes this validation sequence:
- Validate the tool contract — The JSON payload must match the schema generated by
tool.BuiltinContractEntries. - Run the Guardian — The safety reviewer checks for prohibited content in the pending tool call.
- Apply permission rules — The system evaluates whether the call complies with explicit policy constraints.
- Record the evidence — Successful calls append to the session log with full evidence chains.
If any step fails, the host returns an error and rejects the turn, forcing the model to retry or report blockage.
Practical Implementation Examples
Evaluating Output via CLI
When using the interactive REPL, you explicitly close evaluation loops:
# Start a Reasonix session
reasonix
# Execute a command that produces output
> bash "git status"
# Report turn disposition
> update_goal {"status":"continue","reason":"need to add files","next_action":"git add ."}
# Sign off the step with evidence
> complete_step {"step":"Check repository status","result":"untracked files found","evidence":[{"kind":"verification","summary":"git status output"}]}
Programmatic Evaluation with the Go SDK
The Go SDK enforces the same validation layers when evaluating output programmatically:
import (
"context"
"github.com/esengine/DeepSeek-Reasonix/sdk/go"
)
func main() {
client := reasonix.NewClient()
// Execute tool and capture output
out, _ := client.CallTool(context.Background(), reasonix.Tool{
Name: "bash",
Args: map[string]any{
"command": "go test ./...",
},
})
// Evaluate: Report completion status
client.CallTool(context.Background(), reasonix.Tool{
Name: "update_goal",
Args: map[string]any{
"status": "complete",
"reason": "All tests passed",
},
})
// Sign off with evidence
client.CallTool(context.Background(), reasonix.Tool{
Name: "complete_step",
Args: map[string]any{
"step": "Run unit tests",
"result": "passed",
"evidence": []any{
map[string]any{
"kind": "verification",
"summary": out,
},
},
},
})
}
Custom Plugin Evaluation
Custom MCP plugins must also report evidence through the evaluation pipeline:
func MyPluginAction(ctx context.Context, args map[string]any) (any, error) {
result := "generated code"
// Register evidence with Reasonix
_, err := reasonix.CallTool(ctx, reasonix.Tool{
Name: "complete_step",
Args: map[string]any{
"step": "Generate code",
"result": "ok",
"evidence": []any{
map[string]any{
"kind": "manual",
"summary": "Human verified the generated code",
},
},
},
})
return result, err
}
The Guardian and permission engine evaluate these plugin calls identically to built-in tools.
Summary
- Contract validation in
internal/tool/contract.goensures all tool calls match declared JSON schemas. - Status reporting via
update_goaland evidence submission viacomplete_stepare mandatory for finalizing outputs. - Safety review by the Guardian in
internal/guardian/guardian.gofilters prohibited content before acceptance. - Permission checks through
permission.Decideenforce policy compliance based on goal context. - Session logging in
internal/agent/run_loop.gomaintains immutable evidence chains for audit and replay.
Frequently Asked Questions
What happens if I don't call complete_step after a tool execution?
The host will reject the turn and return an error. Reasonix requires either an update_goal report for every turn, and complete_step with at least one evidence item when finishing planned steps. Without this verification, the output remains in a pending state and is not committed to the session log.
Can the Guardian reviewer reject valid tool calls?
Yes. The Guardian in internal/guardian/guardian.go operates independently of schema validation. Even if a tool call matches its JSON contract and passes permission checks, the safety reviewer can reject it for containing unsafe content or violating security policies, forcing the model to revise its approach.
What types of evidence does complete_step accept?
According to internal/tool/builtin/completestep.go, evidence items must include a kind field (such as diff, verification, or manual) and a summary string. The host rejects any complete_step call that lacks at least one evidence item, ensuring all reported results are verifiable.
How does the permission engine handle conflicting rules?
The permission.Decide function in internal/permission/permission.go evaluates rules in the context of the current goal and any explicit "ask" configurations. It applies require, prefer, and ask directives hierarchically, denying execution if mandatory constraints are violated while allowing optional preferences to guide selection when multiple valid paths exist.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →