# Negative Results in Coding Agent Experiments: Why Generated Code Fails

> Discover negative results in AI coding agent experiments. Learn why generated code fails due to sandbox crashes, constraint errors, and off-by-one bugs impacting accuracy and latency.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-23

---

**While the coding agent experiments documented in Chapter 5 of the AI Agent Book demonstrate 40-80% accuracy improvements over pure chain-of-thought reasoning, the negative results include fatal sandbox crashes, ill-posed constraint generations, and propagated off-by-one errors that force the agent to fall back to less accurate baselines or incur significant latency penalties.**

Chapter 5 of the `bojieli/ai-agent-book` repository presents three controlled coding agent experiments (designated 5-1, 5-2, and 5-3) that quantify how augmenting large language models with executable Python affects problem-solving across mathematical, logical, and policy-driven tasks. Despite showing that code assistance dramatically outperforms pure reasoning, the negative results in coding agent experiments reveal critical failure modes rooted in code generation quality and execution environment fragility, as summarized in [`chapter5/README.en.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter5/README.en.md).

## Experiment 5-1: Code-for-Math and Sandbox Instability

According to [`chapter5/code-for-math/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter5/code-for-math/README.md), the first experiment tests a **CodingAgent** on 100 competitive mathematics problems. The baseline pure chain-of-thought (CoT) approach achieves approximately 53% accuracy, while the code-assisted variant that translates problems into executable **sympy**/**numpy** scripts reaches approximately 94% accuracy.

### Runtime Crashes and Fallback Behavior

The primary negative result emerges when the generated code triggers **infinite loops** or **memory-limit exceedances** within the sandbox environment. When the sandbox crashes, the agent falls back to the less accurate pure-CoT path, effectively losing the 41-percentage-point accuracy advantage gained through code execution.

### Syntax Errors and Latency Penalties

A secondary failure mode involves **syntax errors** in the generated Python that abort execution. Rather than producing a valid mathematical result, the agent must trigger a retry loop that adds substantial latency to the inference pipeline without guaranteeing eventual success.

## Experiment 5-2: Code-for-Logic and Constraint Violations

The second experiment, detailed in [`chapter5/code-for-logic/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter5/code-for-logic/README.md), addresses "Knights and Knaves" logic puzzles. Here, the baseline pure-CoT solver attains roughly 48% accuracy, while the code-assisted approach constructs **constraint satisfaction problems (CSPs)** using `python-constraint` to reach approximately 87% accuracy.

### Ill-Posed Constraint Generations

A critical negative result occurs when the agent generates **ill-posed CSPs** with missing variables or contradictory constraints. In these cases, the constraint solver does not return a "no-solution" response; instead, it produces an **incorrect answer** that satisfies the malformed constraints while violating the original puzzle logic.

### Logical Over-Fitting to Generated Rules

The extra reasoning step introduces **over-fitting**: the solver returns a solution that technically satisfies the generated code’s constraints while violating the semantic intent of the natural language puzzle. This creates false confidence in incorrect answers that the pure-CoT baseline would not generate.

## Experiment 5-3: Small-Model Policy Errors

As recorded in [`chapter5/small-model-codified-rules/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter5/small-model-codified-rules/README.md), the third experiment involves a small model (approximately 1B parameters) executing business policy rules. Moving policies from prompt-based instructions to executable Python functions improves task success from approximately 61% to 93%.

### Propagation of Off-by-One Errors

The negative results here stem from **off-by-one errors** and **ambiguous conditionals** in the generated rule code. Unlike syntax errors that crash immediately, these semantic bugs execute successfully, causing the agent to **accept and propagate** the faulty rule downstream, resulting in confident but incorrect decisions.

### Latency Overhead and Rollback Mechanisms

Real-time code validation introduces a **latency penalty** of approximately 1.8× compared to the non-code baseline. Additionally, when the validation system detects rule conflicts, it triggers **rollback** procedures that discard partially completed reasoning chains, further degrading response times without guaranteeing correctness recovery.

## Implementation Architecture and Failure Handling

The core implementation resides in [`chapter5/coding-agent/agents.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter5/coding-agent/agents.py), where the **CodingAgent** class orchestrates the code generation and execution pipeline. The repository provides a sandbox utility (referenced in [`chapter5/coding-agent/sandbox.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter5/coding-agent/sandbox.py)) designed to intercept the negative outcomes described above.

The following example demonstrates the standard execution flow and implicit fallback behavior when the sandbox encounters an error:

```python
from agents import CodingAgent  # see chapter5/coding-agent/agents.py

agent = CodingAgent(model="claude-2")          # a Claude-style LLM

question = "Compute the integral ∫₀¹ x³ dx."

# 1. Generate executable Python code from the math problem

code = agent.run(
    prompt=f"Translate the following math problem into executable Python code using sympy:\n{question}"
)

# 2. Execute safely; crashes trigger fallback to pure CoT

from sandbox import exec_code_safely
result = exec_code_safely(code)

print("Answer:", result)   # → 0.25 if successful, or fallback result if sandbox fails

```

When `exec_code_safely()` encounters the sandbox crashes or syntax errors documented in experiments 5-1 through 5-3, the agent automatically reverts to the pure-CoT reasoning path, inheriting the baseline’s lower accuracy but maintaining system stability.

## Summary

- **Code generation failures** (syntax errors, infinite loops) force the agent to abandon code-assisted reasoning and fall back to less accurate pure-CoT baselines.
- **Constraint solver failures** in logic tasks produce incorrect answers rather than graceful "no-solution" responses when faced with ill-posed problem formulations.
- **Semantic code bugs** in policy rules (off-by-one errors, ambiguous conditionals) execute without crashing, allowing errors to propagate downstream with high confidence.
- **Latency penalties** of approximately 1.8× and rollback procedures in the small-model experiment demonstrate that code assistance introduces non-trivial computational overhead.
- **Sandbox fragility** remains the primary bottleneck: the reliability of code assistance is bounded by the execution environment’s ability to handle malformed or resource-intensive generated code.

## Frequently Asked Questions

### What were the main negative results in the math coding experiment (5-1)?

The primary negative results involved **sandbox crashes** caused by infinite loops or memory limit exceedances, which forced the agent to fall back to the 53% accuracy pure-CoT baseline instead of maintaining the 94% accuracy code-assisted result. Additionally, syntax errors in generated code introduced retry latency without accuracy benefits.

### How did constraint solver failures manifest in the logic experiment (5-2)?

When the agent generated **ill-posed constraint satisfaction problems** with missing variables, the solver did not indicate "no solution" but instead returned answers that satisfied the malformed constraints while violating the original puzzle logic. This created **over-fitting** where the code technically executed correctly but produced semantically wrong answers.

### Why did the small model experiment (5-3) experience latency penalties?

The experiment incurred approximately **1.8× slower response times** because the agent had to execute generated Python functions for policy validation in real-time. When rule conflicts were detected, **rollback mechanisms** further increased latency by discarding and recomputing reasoning steps.

### Do the negative results outweigh the benefits of code assistance?

No. Despite the **negative results in coding agent experiments**, the net accuracy gains remain strongly positive (40-80% improvements across all three tasks). However, the failures demonstrate that code assistance reliability is strictly bounded by **code generation quality** and **sandbox robustness**, necessitating fallback mechanisms for production deployments.