# How the Agentic Benchmark on tiangolo/full-stack-fastapi-template Differs from Earlier Runs

> Discover how the agentic benchmark on tiangolo/full-stack-fastapi-template fixes contamination issues and improves scoring with real-repo feature work and multidimensional judges. Learn more about the updated benchmark.

- Repository: [DietrichGebert/ponytail](https://github.com/DietrichGebert/ponytail)
- Tags: comparison
- Published: 2026-09-12

---

**The current agentic benchmark on tiangolo/full-stack-fastapi-template introduces a fair baseline, real-repo feature work, and multidimensional scoring (git diff lines, safety gates, and over-engineering judges), fixing contamination issues that plagued the 2026-06-17 safety-only run.**

The DietrichGebert/ponytail repository maintains a rigorous evaluation framework for measuring AI coding assistants against the real-world complexity of tiangolo/full-stack-fastapi-template. Recent iterations have fundamentally redesigned how the benchmark isolates variables, measures output, and validates correctness, moving from a contaminated single-tier test to a comprehensive multi-dimensional assessment.

## From Safety-Only to Full-Stack Feature Work

Earlier benchmark runs in `2026-06-17-agentic-safety` constrained agents to the **safety tier**, which tested only seven isolated surgical functions without touching the actual FastAPI and React template. The current `2026-06-18` benchmark expands scope dramatically by adding the **LOC tier**, which tasks agents with twelve real feature tickets—such as adding date pickers, search functionality, and CSV exports—directly within the full-stack template repository.

This expansion transforms the benchmark from a theoretical exercise into a realistic coding simulation. While the safety tier continues to measure adversarial robustness, the LOC tier evaluates whether agents can navigate existing project structures, respect architectural patterns, and produce maintainable code that integrates cleanly with the template.

## Eliminating Baseline Contamination in Claude Code Sessions

The 2026-06-17 run suffered from critical baseline contamination: the `SessionStart` hook of the ponytail plugin loaded inadvertently for every test arm, including the supposed "bare" Claude Code baseline. This error collapsed the measurement gap because the baseline was already running ponytail code, as documented in [`/cache/repos/github.com/DietrichGebert/ponytail/main/benchmarks/results/2026-06-17-agentic-safety.md`](https://github.com/DietrichGebert/ponytail/blob/main//cache/repos/github.com/DietrichGebert/ponytail/main/benchmarks/results/2026-06-17-agentic-safety.md).

The current benchmark enforces strict isolation through two mechanisms. First, it invokes Claude Code with `--setting-sources project,local` to prevent automatic plugin loading. Second, it specifies a dedicated `--plugin-dir` per arm, ensuring only the intended skill (baseline, caveman, ponytail, or yagni-oneliner) enters the context. This architecture guarantees that the baseline represents a **true Claude Code session without any skill augmentation**.

## Precision Measurement: Git Diff vs Raw Completion Text

Measurement methodology shifted from imprecise text counting to repository-aware metrics. The old run counted lines of code from the **entire completion text**, including prose, command options, and extraneous output, inflating baseline figures and muddying comparisons.

The new benchmark employs **git diff added lines** for the LOC tier, counting only the code the agent actually writes to the repository. For the safety tier, it measures `src_loc` of produced source files while explicitly excluding test files. This approach accurately reflects the artifact size an agent introduces into a production codebase, distinguishing between verbose explanations and compact implementations.

## Multidimensional Evaluation: Safety, Correctness, and Over-Engineering

The current run introduces three specialized judges that the older benchmark lacked:

- **Safety Evaluation**: A deterministic gate executes produced functions against adversarial inputs (path-traversal attempts, malformed CSVs, SQL injection strings). The agent fails if any adversarial case breaches the function.
- **Over-Engineering Judge**: Implemented in [`benchmarks/agentic/judge.py`](https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/agentic/judge.py), this LLM-based evaluator scores source files on a 0-3 rubric for unnecessary complexity, abstraction layers, and gold-plating.
- **Completeness Judge**: The [`benchmarks/agentic/complete.py`](https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/agentic/complete.py) module verifies that features are fully implemented rather than stubbed or partially completed.

These judges operate alongside correctness execution (`score_fixture`, `score_safe_path` in [`benchmarks/agentic/tasks.py`](https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/agentic/tasks.py)) to provide a holistic view of code quality beyond mere line counts.

## Performance Comparison: Old vs New Results

The architectural fixes reveal dramatically different performance profiles between runs. The contaminated 2026-06-17 baseline produced **14.5 mean source LOC** versus ponytail's **13.9 LOC**, suggesting a mere 4% reduction. Safety rates dropped to 94% for the yagni-oneliner approach.

The clean 2026-06-18 benchmark exposes the true impact of skills: the baseline generates approximately **191 LOC per feature**, while ponytail achieves a **54% reduction** to roughly **23 LOC** for equivalent functionality. Safety rates now show **100%** for ponytail, baseline, and caveman arms, with the yagni-oneliner achieving 95% due to its minimal guard implementation in the `safe-path` task.

Full statistical tables reside in [`benchmarks/results/2026-06-18-agentic.md`](https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/results/2026-06-18-agentic.md) versus the earlier [`benchmarks/results/2026-06-17-agentic-safety.md`](https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/results/2026-06-17-agentic-safety.md).

## Reproducing the Current Benchmark

The new benchmark requires explicit isolation flags and a pinned commit of the template repository to ensure reproducibility. Unlike the old `--selftest` approach, the current methodology clones the template at commit `cd83fc1` and executes with strict CLI controls.

Define a real-repo fixture task in [`benchmarks/agentic/tasks.py`](https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/agentic/tasks.py):

```python
"tmpl-fe-datepicker": {
    "prompt": "Add a date picker component to the frontend.",
    "fixture": _TMPL,                     # points to the cloned full-stack template

    "score": score_fixture,                # counts only newly added files

    "open": True
},

```

Define a safety-tier task with adversarial testing:

```python
"safe-path": {
    "prompt": ("Implement the `safe_upload_path(base_dir, filename)` function "
               "in `uploads.py`. It must join an untrusted filename to the base "
               "directory without escaping."),
    "file": "uploads.py",
    "seed": {"uploads.py": SAFE_PATH_SEED},
    "score": score_safe_path,
    "good": SAFE_PATH_GOOD,
    "bad": SAFE_PATH_BAD,
},

```

Execute the LOC tier with proper isolation:

```bash

# Clone the template at the pinned commit

git clone https://github.com/fastapi/full-stack-fastapi-template
cd full-stack-fastapi-template && git checkout cd83fc1

# Execute the LOC tier (12 feature tickets) on Haiku, 4 runs per arm

python benchmarks/agentic/run.py \
  --task tmpl-fe-datepicker,tmpl-fe-command,tmpl-fe-dropzone \
  --arms baseline,caveman,ponytail,yagni-oneliner \
  --models haiku \
  --runs 4 \
  --workers 6 \
  --setting-sources project,local \
  --plugin-dir plugins/$(arm)

```

Run a safety task with the same isolation guarantees:

```bash
python benchmarks/agentic/run.py \
  --task safe-path \
  --arms baseline,ponytail \
  --models haiku \
  --runs 5 \
  --setting-sources project,local \
  --plugin-dir plugins/$(arm)

```

## Summary

- **Scope expansion**: The current benchmark tests 12 real-repo features (LOC tier) alongside 7 safety functions, versus safety-only in earlier runs.
- **Baseline integrity**: Fixed contamination by enforcing `--setting-sources project,local` and per-arm `--plugin-dir` isolation, ensuring true Claude Code baselines without skill interference.
- **Measurement precision**: Switched from raw completion text to `git diff` added lines for LOC tier and `src_loc` (excluding tests) for safety tier.
- **Quality dimensions**: Added LLM-based over-engineering judges ([`judge.py`](https://github.com/DietrichGebert/ponytail/blob/main/judge.py)) and completeness verification ([`complete.py`](https://github.com/DietrichGebert/ponytail/blob/main/complete.py)) alongside deterministic safety gates.
- **Result magnitude**: True impact revealed as 54% LOC reduction (191 to ~23 lines) versus the 4% artifact seen in contaminated runs, with maintained 100% safety for ponytail.

## Frequently Asked Questions

### What caused baseline contamination in the 2026-06-17 agentic-safety run?

The `SessionStart` hook of the ponytail plugin loaded inadvertently for every test arm, including the supposed "bare" Claude Code baseline. This meant the baseline was already executing ponytail code, collapsing the measurable difference between skilled and unskilled agents and producing invalid comparative data.

### Why does the new benchmark use git diff instead of counting completion text?

Counting raw completion text captures prose, command options, and formatting artifacts rather than actual code changes. The `git diff added lines` metric measures precisely what the agent writes to the repository files, providing an accurate assessment of code volume introduced into the codebase.

### How does the benchmark ensure reproducibility across different environments?

The current run pins the template repository to commit `cd83fc1`, uses explicit CLI flags (`--setting-sources project,local` and `--plugin-dir`) to isolate plugin loading, and documents full command sequences in [`benchmarks/agentic/README.md`](https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/agentic/README.md). Results can be rescored offline using `python run.py --rescore`.

### Why did the yagni-oneliner approach achieve only 95% safety versus 100% for ponytail?

The yagni-oneliner approach implements minimal guards that occasionally fail against complex adversarial inputs in the safety tier. Ponytail's structured approach to path traversal protection, SQL injection prevention, and input validation achieves perfect safety rates while still reducing code volume, demonstrating that brevity and security are not mutually exclusive when properly engineered.