DeepSeek‑Reasonix Performance Benchmarks: Micro‑Benchmarks for Extension Runtime Latency

DeepSeek‑Reasonix provides a suite of micro‑benchmarks targeting sub‑50 ms latency thresholds for core extension runtime operations including dependency graph construction, runtime plan diffing, and side‑car process management. These benchmarks are designed as "soft CI" guards that catch order‑of‑magnitude regressions without flaky failures.

The performance benchmark for DeepSeek‑Reasonix focuses exclusively on the extension runtime—the subsystem responsible for building dependency graphs, computing runtime plans, and orchestrating side‑car processes. This article breaks down the specific metrics, thresholds, and implementation details derived directly from the source code in esengine/DeepSeek-Reasonix.

Core Runtime Benchmarks and Thresholds

The repository defines explicit latency targets for three critical operations in internal/extension/bench_threshold_test.go. These thresholds represent soft‑fail limits on typical developer hardware.

Metric Scenario Soft‑CI Target Source File
BuildDependencyGraph 32 components < 50 ms bench_threshold_test.go
DiffRuntimePlan (no‑op) Identical graph comparison < 20 ms bench_threshold_test.go
EffectScope.Dispose 64 effects cleanup < 50 ms bench_threshold_test.go

These thresholds are deliberately generous. They serve as regression detectors rather than strict gates, ensuring CI stability while flagging performance degradations that exceed an order of magnitude.

Extension Kernel and Graph Scaling Benchmarks

Beyond fixed thresholds, the repository includes statistical micro‑benchmarks in internal/extension/benchmark_test.go that report p50 and p95 latencies across varying workload sizes.

Extension Kernel Startup

The BenchmarkExtensionKernelStartup function measures the cost of building a RuntimeSnapshot from a catalog of contributions:

  • Empty snapshot (no extensions): p50 ≈ 0.9 ms
  • 64 interceptors: p95 ≈ 1.5 ms

These figures are hardware‑dependent and represent the immutable snapshot construction performed by internal/extension/builder.go.

Dependency Graph and Plan Scaling

BenchmarkDependencyGraphAndPlan exercises the resolver across three graph sizes:


# Run graph and plan benchmarks with memory profiling

go test ./internal/extension/ \
  -bench 'BenchmarkDependencyGraphAndPlan' \
  -benchmem -count=3

The benchmark reports latency per component count for:

  • Graph construction (8, 64, 256 components)
  • Plan diff operations (no‑op and full transition cases)

The no‑op diff validates that comparing identical graphs remains cheap, while the full diff stresses the planner with version bumps across all components.

Side‑Car Process Benchmarks

Side‑cars in DeepSeek‑Reasonix run as separate processes communicating via NDJSON RPC over the Extension Protocol. The benchmarks in internal/extension/sidecar/benchmark_test.go isolate two critical latencies.

Side‑Car Startup Overhead

BenchmarkExtensionSidecarStartup measures the complete handshake lifecycle:

Percentile Latency
p50 ≈ 1.2 ms
p95 ≈ 2.4 ms

This includes process spawning and the StartClient handshake protocol.

Dispatch Latency Scaling

BenchmarkExtensionSidecarDispatchLatency characterizes turn path and tool path RPC overhead:

  • 1 side‑car: Baseline dispatch latency
  • 4 side‑cars: Linear scaling verification

Latency grows roughly linearly with side‑car count, allowing developers to extrapolate overhead for multi‑side‑car configurations.


# Run dispatch latency with increased sample count for stable p95

go test ./internal/extension/sidecar/ \
  -bench BenchmarkExtensionSidecarDispatchLatency \
  -count=5

Running the Benchmark Suite

Execute the complete DeepSeek‑Reasonix performance benchmark suite using Go's standard testing harness:


# Soft-CI threshold tests (single run, fast feedback)

go test ./internal/extension/ \
  -run 'TestGraphAndPlanLatencyBaseline|TestEffectScopeDisposeBaseline' \
  -count=1

# Full micro-benchmarks with statistical stability (3 repeats)

go test ./internal/extension/ \
  -bench 'BenchmarkDependencyGraphAndPlan|BenchmarkExtensionKernelStartup' \
  -benchmem -count=3

Implementation Targets

Understanding what each benchmark exercises requires familiarity with the underlying architecture:

  • BuildDependencyGraph (internal/extension/graph.go): Resolves ComponentDescriptor slices into directed acyclic graphs
  • DiffRuntimePlan: Computes minimal transition plans between graph states; critical for incremental updates
  • EffectScope (internal/extension/): Tracks reversible resources; the dispose benchmark ensures bulk cleanup performance
  • Side‑car layer (internal/extension/sidecar/): Process isolation for extension code; benchmarks validate IPC efficiency

Example: Measuring Graph Construction

// Reproduce the BuildDependencyGraph benchmark locally
func ExampleGraphBuild() {
    comps := make([]ComponentDescriptor, 32)
    for i := range comps {
        comps[i] = ComponentDescriptor{ID: ComponentID(fmt.Sprintf("c%d", i))}
    }
    
    start := time.Now()
    g, err := BuildDependencyGraph(comps)
    if err != nil {
        log.Fatal(err)
    }
    elapsed := time.Since(start)
    
    fmt.Printf("graph built in %v (target: <50ms)\n", elapsed)
    // Use graph...
    _ = g
}

Key Source Files

File Purpose
internal/extension/benchmark_test.go Kernel startup, graph building, plan diffing micro‑benchmarks
internal/extension/bench_threshold_test.go Soft‑CI threshold assertions for core operations
internal/extension/sidecar/benchmark_test.go Side‑car startup and dispatch latency benchmarks
internal/extension/builder.go RuntimeSnapshot construction implementation
internal/extension/graph.go Dependency graph resolver and plan diffing
docs/EXTENSION_RUNTIME_V2_PERF.md Human‑readable performance documentation

Summary

  • DeepSeek‑Reasonix performance benchmarks target sub‑50 ms latency for graph construction and effect disposal, with sub‑20 ms for plan diffing
  • Statistical benchmarks report p50/p95 percentiles rather than single‑point measurements
  • Side‑car benchmarks validate process startup (~1–2 ms) and linear dispatch scaling
  • All benchmarks implement soft CI thresholds to prevent flaky failures while catching regressions
  • Source files in internal/extension/ provide both the benchmark harness and the measured implementation

Frequently Asked Questions

What hardware do the DeepSeek‑Reasonix benchmark targets assume?

The soft‑CI thresholds assume a typical developer machine—not dedicated performance hardware. Targets are intentionally conservative (sub‑50 ms for 32–64 component operations) to accommodate variable hardware while still detecting order‑of‑magnitude performance regressions. Production deployments on faster hardware will observe significantly better latencies.

How does DiffRuntimePlan maintain sub‑20 ms performance on identical graphs?

The no‑op case leverages structural sharing in the immutable RuntimeSnapshot. When graphs are identical, the diff algorithm short‑circuits by comparing root pointers before traversing component descriptors. This optimization is validated by TestGraphAndPlanLatencyBaseline in bench_threshold_test.go.

Why are side‑car benchmarks separated from core extension benchmarks?

Side‑cars introduce process boundary overhead and OS‑level variability (process scheduling, IPC buffer management) that would distort measurements of in‑process operations. Isolating them in internal/extension/sidecar/benchmark_test.go allows targeted optimization of the NDJSON RPC layer without confounding core runtime metrics.

Can I run individual benchmarks without the full suite?

Yes. The Go benchmark naming convention enables precise filtering:


# Run only the 256‑component graph benchmark

go test ./internal/extension/ -bench BenchmarkDependencyGraphAndPlan/256 -count=3

# Run only side‑car startup (not dispatch)

go test ./internal/extension/sidecar/ -bench BenchmarkExtensionSidecarStartup$

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →