Methodology Behind the Intended-vs-Implemented Audit Skill: A 5-Step Security Framework

The intended-vs-implemented audit skill applies a rigorous five-step methodology to compare documented system boundaries against actual code enforcement, producing actionable security findings only when discrepancies cross trust, cost, data, or tenant boundaries.

The intended-vs-implemented skill in the phuryn/pm-skills repository provides a systematic framework for auditing gaps between documented security requirements and their implementation in source code. Unlike generic code reviews, this methodology treats documentation as untrusted input that must be verified against concrete enforcement points, ensuring that every reported finding includes specific file references, attacker scenarios, and remediation steps.

The 5-Step Audit Methodology

According to the skill definition in pm-ai-shipping/skills/intended-vs-implemented/SKILL.md, the audit proceeds through five discrete phases:

1. Establish Documented Intent

The audit begins by reading the markdown documentation set—typically files like permissions.md and architecture.md within a documentation/ directory—to establish the source of truth for intended security and functional boundaries. These documents define the rules, scopes, and classifications that the system is supposed to enforce, such as "admin-only endpoint" or "tenant-isolated data."

2. Gather Implementation Evidence

The methodology requires locating the concrete enforcement points in the codebase that correspond to each documented rule. This includes authorization checks, query filters, and input sanitizers. Crucially, evidence must be cited as precise file and line references—vague comments or assumptions are insufficient. As implemented in phuryn/pm-skills, this step demands provable presence or absence of enforcement mechanisms.

3. Compare Claims to Evidence

For every documented rule, the audit verifies whether an enforcement point exists on every execution path. The comparison happens one boundary at a time, checking that the documented intent aligns with the implemented reality. Inconsistent or missing checks are flagged as potential mismatches for further classification.

4. Classify Mismatches by Security Impact

Not every discrepancy warrants a finding. The methodology classifies mismatches based on whether they allow an unauthorized actor to reach sensitive data, financial resources, infrastructure, or another tenant’s environment. Cosmetic drifts—where documentation and code differ but no security boundary is crossed—are explicitly ignored.

5. Avoid Hand-Wavy Findings

Every reported gap must include four specific elements: the exact documented intent, the exact code location, the attacker/victim scenario, and a concrete fix. If any of these elements cannot be cited with precision, the item is treated as an open investigation rather than a definitive finding. This prevents the generation of unactionable audit noise.

Defining Intent, Evidence, and Meaningful Mismatches

The methodology establishes strict definitions for three critical concepts:

  • Intent — A documented rule, boundary, scope, or classification (e.g., "admin-only endpoint") sourced from the documentation/*.md files.

  • Implementation Evidence — A cited enforcement point in the source code (or provable absence thereof), referenced by specific file path and line number.

  • Meaningful Mismatch — A discrepancy between documented rules and code that crosses a trust, cost, data, or tenant boundary. Only these qualify as security findings; cosmetic inconsistencies are filtered out.

Running the Audit: Practical Examples

The skill integrates into the shipping workflow through specific commands and can be invoked manually for custom CI pipelines.

Running the Static Security Audit

To execute the intended-vs-implemented audit as part of a security review:


# From the repository root, trigger the static security audit

# which internally applies the intended-vs-implemented skill

$ /security-audit-static

This command initiates the five-step audit process described in the shipping sequence, analyzing documentation against the source tree.

Full Shipping Check Integration

For a comprehensive review that includes the audit alongside performance and coverage checks:


# Generate a complete shipping packet including the audit

$ /ship-check

The /ship-check command orchestrates the full audit sequence, combining documentation validation, security auditing using this skill, and test-coverage mapping as defined in pm-ai-shipping/commands/ship-check.md.

Programmatic Usage in Custom Scripts

Developers can embed the logic directly into Python-based CI pipelines:


# Re-use the skill logic in custom automation

from pm_ai_shipping.skills.intended_vs_implemented import audit_intent_vs_code

docs = load_markdown_folder('documentation/')    # Intent sources

code = load_source_tree('src/')                  # Implementation evidence

findings = audit_intent_vs_code(docs, code)

for finding in findings:
    print(f"Intent: {finding.intent}")
    print(f"Location: {finding.code_location}")
    print(f"Risk: {finding.risk}, Fix: {finding.fix}")

The audit_intent_vs_code function encapsulates the complete methodology, returning structured findings that include the risk classification and concrete remediation steps.

Key Source Files and Implementation

The methodology is implemented across several files in the phuryn/pm-skills repository:

Summary

  • The intended-vs-implemented skill uses a five-step methodology: establish intent, gather evidence, compare claims, classify mismatches, and require concrete citations.

  • Documented-but-unenforced rules are always reported as findings ranked by risk, while undocumented-but-enforced rules are flagged as documentation staleness, not security issues.

  • Only mismatches that cross trust, cost, data, or tenant boundaries qualify as meaningful security findings; cosmetic drifts are ignored.

  • Every finding must cite the exact documentation, exact code location, attacker scenario, and concrete fix—otherwise it remains an open investigation.

  • The methodology treats both documentation and code as untrusted inputs, never fabricating intent when documentation is missing.

Frequently Asked Questions

What is the difference between documented-but-unenforced and undocumented-but-enforced rules?

Documented-but-unenforced rules occur when the documentation specifies a security boundary (such as "only admins can access this endpoint") but the code lacks the corresponding enforcement check—these are always reported as security findings. Undocumented-but-enforced rules occur when the code implements a restriction that isn't mentioned in the documentation—these are flagged as stale or incomplete documentation rather than security vulnerabilities, since the system is actually protected.

How does the audit handle missing documentation?

If no documentation exists for a particular component, the methodology explicitly refuses to fabricate intent. Instead of assuming what the system should do, the audit reports that the documentation is missing. This prevents the introduction of false assumptions about security requirements and forces teams to formally document boundaries before they can be verified.

What makes a mismatch "meaningful" versus cosmetic?

A mismatch is considered meaningful only when the gap between documentation and implementation allows an unauthorized actor to access sensitive data, incur costs, touch infrastructure, or cross into another tenant's environment. If the discrepancy involves internal refactoring, naming conventions, or architectural patterns that do not affect security boundaries, it is classified as cosmetic and excluded from findings.

Can the intended-vs-implemented skill be used outside of security audits?

Yes. While the methodology feeds into security and performance audits, the underlying audit_intent_vs_code function can be imported and used in any CI pipeline or custom script to verify that functional requirements documented in markdown files match the actual code implementation. It does not replace lower-level static analysis but serves as a high-level consistency check between specification and reality.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →