# How to Check for Gaps Between AI Code Intent and Implementation

> Discover how to check for gaps between AI code intent and implementation. The PM-AI-Shipping plugin reveals security-critical mismatches missed by static analyzers.

- Repository: [Pawel Huryn/pm-skills](https://github.com/phuryn/pm-skills)
- Tags: how-to-guide
- Published: 2026-07-09

---

**The PM-AI-Shipping plugin provides a systematic three-phase workflow to surface mismatches between documented intent and actual code implementation, exposing security-critical gaps that static analyzers miss.**

The phuryn/pm-skills repository delivers specialized tools for shipping AI-generated code safely. When developing applications with AI assistants, the generated code often diverges from the intended behavior, creating silent security vulnerabilities. Checking for gaps between AI code intent and implementation requires comparing explicit markdown documentation against concrete enforcement points in the codebase.

## The Three-Phase Gap Detection Workflow

### Phase 1: Capture Intent with Shipping Artifacts

The **shipping-artifacts skill** defined in [`pm-ai-shipping/skills/shipping-artifacts/SKILL.md`](https://github.com/phuryn/pm-skills/blob/main/pm-ai-shipping/skills/shipping-artifacts/SKILL.md) establishes the documentation set that serves as the source of truth for intended behavior.

Core documentation includes:
- **Architecture** specifications and **data flows**
- **Permission models** and **business rules**
- **Environment variables** and **secrets handling**
- **Test specifications** and validation criteria

Conditional documentation covers optional capabilities like email automation, cron jobs, SEO requirements, and third-party integrations. These are only generated when the corresponding code capabilities exist.

### Phase 2: Audit Implementation vs Intent

The **intended-vs-implemented skill** located in [`pm-ai-shipping/skills/intended-vs-implemented/SKILL.md`](https://github.com/phuryn/pm-skills/blob/main/pm-ai-shipping/skills/intended-vs-implemented/SKILL.md) performs the actual gap analysis.

This audit process:
1. Reads the markdown documentation files as the canonical source of truth
2. Traverses the codebase to locate concrete enforcement points (authentication checks, authorization filters, input sanitizers)
3. Identifies **mismatches that matter**—divergences crossing trust boundaries, cost boundaries, data boundaries, or tenant boundaries

Unlike generic static analyzers that verify internal consistency, this comparison exposes high-value bugs such as permissions documented but never enforced, or public endpoints skipping expected denial cases.

### Phase 3: Report Security-Critical Findings

The audit generates a structured report containing:
- The exact **documented claim** quoted from the markdown file
- The **implementation evidence** with precise file path and line number
- **Attacker and victim personas** enabled by the gap
- Concrete **remediation suggestions**

## Step-by-Step Workflow

Install the plugin and execute the gap detection pipeline:

1. **Install the PM-AI-Shipping plugin**:

```bash
claude plugin marketplace add phuryn/pm-skills
claude plugin install pm-ai-shipping@pm-skills

```

2. **Generate documentation** from your existing codebase:

```bash
cd <repository-root>
claude /document-app

```

This creates [`architecture.md`](https://github.com/phuryn/pm-skills/blob/main/architecture.md), [`flows.md`](https://github.com/phuryn/pm-skills/blob/main/flows.md), [`permissions.md`](https://github.com/phuryn/pm-skills/blob/main/permissions.md), and other core files by reverse-engineering the code.

3. **Derive test coverage mapping**:

```bash
claude /derive-tests

```

This generates [`tests.md`](https://github.com/phuryn/pm-skills/blob/main/tests.md) linking each documented rule to existing test coverage, highlighting uncovered gaps.

4. **Run the intent-implementation audit**:

```bash
claude /intended-vs-implemented

```

5. **Bundle for review** (optional):

```bash
claude /ship-check

```

The `ship-check` command automatically runs the intent-vs-implementation audit as part of its reviewer-ready shipping packet.

### Sample Output

A typical mismatch report identifies the specific divergence:

```text
🔎 MISMATCH #1
📄 Documented intent (permissions.md):
   "Only admins may delete a project."
📁 Code evidence (src/api/projects.ts:112):
   if (user.role !== 'admin') { /* no explicit denial */ }
👤 Attacker: any authenticated user
🛡️ Risk: unauthorized project deletion → data loss
✅ Fix: Add explicit permission check and deny response.

```

## Key Files in the Source Repository

The **phuryn/pm-skills** repository contains the following critical components:

- **[`pm-ai-shipping/skills/shipping-artifacts/SKILL.md`](https://github.com/phuryn/pm-skills/blob/main/pm-ai-shipping/skills/shipping-artifacts/SKILL.md)** – Defines the documentation set structure required to capture intent, distinguishing between core and conditional artifacts.

- **[`pm-ai-shipping/skills/intended-vs-implemented/SKILL.md`](https://github.com/phuryn/pm-skills/blob/main/pm-ai-shipping/skills/intended-vs-implemented/SKILL.md)** – Describes the gap-analysis methodology, classification rules for boundary-crossing mismatches, and the comparison algorithm.

- **[`pm-ai-shipping/commands/ship-check.md`](https://github.com/phuryn/pm-skills/blob/main/pm-ai-shipping/commands/ship-check.md)** – Orchestrates the complete pipeline, bundling documentation, audit reports, and test-coverage maps into a reviewer-ready shipping packet.

- **[`pm-ai-shipping/commands/document-app.md`](https://github.com/phuryn/pm-skills/blob/main/pm-ai-shipping/commands/document-app.md)** – Implements the reverse-engineering logic that generates core documentation from existing source code.

- **[`pm-ai-shipping/commands/derive-tests.md`](https://github.com/phuryn/pm-skills/blob/main/pm-ai-shipping/commands/derive-tests.md)** – Generates the test coverage mapping that links documented rules to existing test implementations.

## Summary

- **Document intent explicitly** using the shipping-artifacts skill to create a source of truth that generic static analyzers lack.
- **Compare documentation against code** using the intended-vs-implemented skill to find mismatches crossing security boundaries.
- **Receive actionable reports** with file paths, line numbers, and remediation steps for each identified gap.
- **Automate the workflow** via the `ship-check` command to continuously validate AI-generated code against its documented requirements.

## Frequently Asked Questions

### What makes intent-implementation gaps different from regular bugs?

Regular bugs represent internal inconsistencies or syntax errors, while intent-implementation gaps occur when the code functionally works but fails to enforce documented security requirements. According to the `intended-vs-implemented` skill documentation, these gaps specifically cross trust, cost, data, or tenant boundaries, enabling attackers to bypass documented permissions or access controls.

### Can I use this workflow on existing codebases that weren't AI-generated?

Yes. The `claude /document-app` command reverse-engineers existing repositories to generate the required documentation set. This makes the workflow applicable to legacy applications, human-written code, or vibe-coded prototypes that need security validation.

### How does the ship-check command differ from running audits individually?

The `ship-check` command defined in [`pm-ai-shipping/commands/ship-check.md`](https://github.com/phuryn/pm-skills/blob/main/pm-ai-shipping/commands/ship-check.md) bundles the entire validation pipeline into a single execution. It automatically runs the intent-vs-implementation audit after generating documentation and test mappings, producing a consolidated shipping packet for reviewers. You can run individual commands like `/intended-vs-implemented` for faster iteration, but `ship-check` ensures comprehensive coverage.

### What types of mismatches does the audit prioritize?

The audit prioritizes **high-value bugs** that typical linters miss, including permissions documented but never checked, missing authorization filters on public endpoints, and incomplete denial cases in conditional logic. The system specifically flags mismatches that enable attackers to exploit trust boundaries or access sensitive data across tenant isolation.