Negative Results in Coding Agent Experiments: Why Generated Code Fails

While the coding agent experiments documented in Chapter 5 of the AI Agent Book demonstrate 40-80% accuracy improvements over pure chain-of-thought reasoning, the negative results include fatal sandbox crashes, ill-posed constraint generations, and propagated off-by-one errors that force the agent to fall back to less accurate baselines or incur significant latency penalties.

Chapter 5 of the bojieli/ai-agent-book repository presents three controlled coding agent experiments (designated 5-1, 5-2, and 5-3) that quantify how augmenting large language models with executable Python affects problem-solving across mathematical, logical, and policy-driven tasks. Despite showing that code assistance dramatically outperforms pure reasoning, the negative results in coding agent experiments reveal critical failure modes rooted in code generation quality and execution environment fragility, as summarized in chapter5/README.en.md.

Experiment 5-1: Code-for-Math and Sandbox Instability

According to chapter5/code-for-math/README.md, the first experiment tests a CodingAgent on 100 competitive mathematics problems. The baseline pure chain-of-thought (CoT) approach achieves approximately 53% accuracy, while the code-assisted variant that translates problems into executable sympy/numpy scripts reaches approximately 94% accuracy.

Runtime Crashes and Fallback Behavior

The primary negative result emerges when the generated code triggers infinite loops or memory-limit exceedances within the sandbox environment. When the sandbox crashes, the agent falls back to the less accurate pure-CoT path, effectively losing the 41-percentage-point accuracy advantage gained through code execution.

Syntax Errors and Latency Penalties

A secondary failure mode involves syntax errors in the generated Python that abort execution. Rather than producing a valid mathematical result, the agent must trigger a retry loop that adds substantial latency to the inference pipeline without guaranteeing eventual success.

Experiment 5-2: Code-for-Logic and Constraint Violations

The second experiment, detailed in chapter5/code-for-logic/README.md, addresses "Knights and Knaves" logic puzzles. Here, the baseline pure-CoT solver attains roughly 48% accuracy, while the code-assisted approach constructs constraint satisfaction problems (CSPs) using python-constraint to reach approximately 87% accuracy.

Ill-Posed Constraint Generations

A critical negative result occurs when the agent generates ill-posed CSPs with missing variables or contradictory constraints. In these cases, the constraint solver does not return a "no-solution" response; instead, it produces an incorrect answer that satisfies the malformed constraints while violating the original puzzle logic.

Logical Over-Fitting to Generated Rules

The extra reasoning step introduces over-fitting: the solver returns a solution that technically satisfies the generated code’s constraints while violating the semantic intent of the natural language puzzle. This creates false confidence in incorrect answers that the pure-CoT baseline would not generate.

Experiment 5-3: Small-Model Policy Errors

As recorded in chapter5/small-model-codified-rules/README.md, the third experiment involves a small model (approximately 1B parameters) executing business policy rules. Moving policies from prompt-based instructions to executable Python functions improves task success from approximately 61% to 93%.

Propagation of Off-by-One Errors

The negative results here stem from off-by-one errors and ambiguous conditionals in the generated rule code. Unlike syntax errors that crash immediately, these semantic bugs execute successfully, causing the agent to accept and propagate the faulty rule downstream, resulting in confident but incorrect decisions.

Latency Overhead and Rollback Mechanisms

Real-time code validation introduces a latency penalty of approximately 1.8× compared to the non-code baseline. Additionally, when the validation system detects rule conflicts, it triggers rollback procedures that discard partially completed reasoning chains, further degrading response times without guaranteeing correctness recovery.

Implementation Architecture and Failure Handling

The core implementation resides in chapter5/coding-agent/agents.py, where the CodingAgent class orchestrates the code generation and execution pipeline. The repository provides a sandbox utility (referenced in chapter5/coding-agent/sandbox.py) designed to intercept the negative outcomes described above.

The following example demonstrates the standard execution flow and implicit fallback behavior when the sandbox encounters an error:

from agents import CodingAgent  # see chapter5/coding-agent/agents.py

agent = CodingAgent(model="claude-2")          # a Claude-style LLM

question = "Compute the integral ∫₀¹ x³ dx."

# 1. Generate executable Python code from the math problem

code = agent.run(
    prompt=f"Translate the following math problem into executable Python code using sympy:\n{question}"
)

# 2. Execute safely; crashes trigger fallback to pure CoT

from sandbox import exec_code_safely
result = exec_code_safely(code)

print("Answer:", result)   # → 0.25 if successful, or fallback result if sandbox fails

When exec_code_safely() encounters the sandbox crashes or syntax errors documented in experiments 5-1 through 5-3, the agent automatically reverts to the pure-CoT reasoning path, inheriting the baseline’s lower accuracy but maintaining system stability.

Summary

  • Code generation failures (syntax errors, infinite loops) force the agent to abandon code-assisted reasoning and fall back to less accurate pure-CoT baselines.
  • Constraint solver failures in logic tasks produce incorrect answers rather than graceful "no-solution" responses when faced with ill-posed problem formulations.
  • Semantic code bugs in policy rules (off-by-one errors, ambiguous conditionals) execute without crashing, allowing errors to propagate downstream with high confidence.
  • Latency penalties of approximately 1.8× and rollback procedures in the small-model experiment demonstrate that code assistance introduces non-trivial computational overhead.
  • Sandbox fragility remains the primary bottleneck: the reliability of code assistance is bounded by the execution environment’s ability to handle malformed or resource-intensive generated code.

Frequently Asked Questions

What were the main negative results in the math coding experiment (5-1)?

The primary negative results involved sandbox crashes caused by infinite loops or memory limit exceedances, which forced the agent to fall back to the 53% accuracy pure-CoT baseline instead of maintaining the 94% accuracy code-assisted result. Additionally, syntax errors in generated code introduced retry latency without accuracy benefits.

How did constraint solver failures manifest in the logic experiment (5-2)?

When the agent generated ill-posed constraint satisfaction problems with missing variables, the solver did not indicate "no solution" but instead returned answers that satisfied the malformed constraints while violating the original puzzle logic. This created over-fitting where the code technically executed correctly but produced semantically wrong answers.

Why did the small model experiment (5-3) experience latency penalties?

The experiment incurred approximately 1.8× slower response times because the agent had to execute generated Python functions for policy validation in real-time. When rule conflicts were detected, rollback mechanisms further increased latency by discarding and recomputing reasoning steps.

Do the negative results outweigh the benefits of code assistance?

No. Despite the negative results in coding agent experiments, the net accuracy gains remain strongly positive (40-80% improvements across all three tasks). However, the failures demonstrate that code assistance reliability is strictly bounded by code generation quality and sandbox robustness, necessitating fallback mechanisms for production deployments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →