How to Test AI Agent Interactions with the emilkowalski/skills Library

The emilkowalski/skills library enables deterministic testing of AI agent interactions by wrapping declarative markdown skill files into a CLI that produces predictable outputs, allowing you to verify gate enforcement, step execution, and result consistency through standard unit and integration testing tools.

The emilkowalski/skills repository provides a declarative framework for AI agents, where each skill is defined as a plain-text markdown file containing hard rules and step-by-step workflows. Because the library exposes these skills through a command-line interface, testing AI agent interactions reduces to invoking the CLI with specific requests and asserting on the generated outputs.

Understanding the Skills Architecture

Each skill in the repository is a self-contained markdown file that describes what an AI agent should do and how it should reason about the request. In skills/animate/SKILL.md, the file structure includes a metadata header with name and description fields, followed by a step-by-step workflow that acts as an automatic gate for the agent.

When a model (such as GPT-4, Claude, or Llama) is wrapped by the skills CLI via npx skills@latest, it receives the full markdown content of the chosen skill as its prompt. The model then parses the hard rules, executes the ordered steps, and returns a concise result that the CLI prints as generated code and decision summaries.

Key files that define this architecture include:

Testing Strategy for AI Agent Interactions

Because the workflow is fully specified in markdown, testing AI agent interactions follows a deterministic validation pattern. You can verify behavior across six critical phases:

Prompt Delivery Verification

Confirm that the model receives the exact markdown of the requested skill. Run npx skills@latest show <skill> and compare the output to the source file content using git diff to ensure no corruption occurs during transmission.

Gate Enforcement Testing

Hard-rule gates automatically reject disallowed requests. To test this, provide a prohibited request (such as "animate a keyboard shortcut") and assert that the CLI prints a gate-rejection message rather than proceeding with code generation.

Step Execution Validation

Each step must produce required artifacts like decision tables or code snippets. Parse the CLI JSON output (using the --json flag if available) and assert the presence of specific fields such as purpose, tool, and properties to verify the agent followed the ordered workflow correctly.

Result Consistency Checks

Feed the generated code to your build pipeline (for example, npm run build) and ensure no lint or type errors arise. This confirms that the AI-produced code compiles and renders with the chosen tool as intended.

Reduced-Motion Handling

For accessibility compliance, verify that media-query gating is included when appropriate. Search the generated CSS or JavaScript for prefers-reduced-motion blocks to ensure the skill automatically handles motion preferences.

Cross-Skill Interaction Testing

When skills delegate to one another (such as animate invoking pick-ui-library), verify correct hand-off by running compound commands and checking that the second skill receives expected context from the first.

Practical Testing Examples

The purely declarative nature of the library means your test harness can be written in any language capable of invoking the CLI, capturing stdout, and comparing it against expectations.

CLI Output Verification with Bash

Test that gates pass and specific CSS properties are generated:


# Install the latest skill package

npx skills@latest add emilkowalski/skills

# Execute the animate skill on a sample request

OUTPUT=$(npx skills@latest run animate \
  --request "Add a fade-in animation to a success toast")

# Check that the gate did not reject the request

echo "$OUTPUT" | grep -q "Gate result" && echo "✅ gate passed"

# Extract the generated CSS and verify it contains the expected easing

echo "$OUTPUT" | grep -q "--ease-out:" && echo "✅ easing present"

Unit Testing with Jest

Verify that the pick-ui-library skill selects trusted dependencies:

import { execSync } from "child_process";

test("pick-ui-library selects a trusted library", () => {
  const result = execSync(
    "npx skills@latest run pick-ui-library --request " +
    "\"Need a toast component for a React app\"",
    { encoding: "utf-8" }
  );

  // The skill should recommend a library that already exists in package.json
  expect(result).toMatch(/react-toastify|react-hot-toast/);
  
  // It must not suggest installing an untrusted package
  expect(result).not.toMatch(/npm i .*dangerous/);
});

End-to-End Browser Testing

Validate that generated animations respect accessibility requirements using Playwright:


# Generate animation code

npx skills@latest run animate \
  --request "Slide in a drawer from the right" > drawer.css

# Append a reduced-motion override (the skill does this automatically)

cat <<'EOF' >> drawer.css
@media (prefers-reduced-motion: reduce) {
  .drawer { animation: none; transform: translateX(0); }
}
EOF

# Run a headless browser test

npx playwright test drawer.spec.ts

Key Reference Files

Understanding these source files is essential for writing comprehensive tests:

Summary

  • Declarative markdown skills in emilkowalski/skills provide deterministic, testable prompts for AI agents
  • Gate enforcement can be verified by submitting prohibited requests and asserting rejection messages
  • Step execution is validated by parsing CLI output for required fields like purpose and tool
  • Build integration confirms that generated code compiles without errors in your target environment
  • Any testing framework can invoke the CLI via npx skills@latest run, making the library language-agnostic

Frequently Asked Questions

What makes skills testable compared to other AI frameworks?

Unlike black-box AI interactions, the emilkowalski/skills library uses declarative markdown files that serve as explicit contracts. The CLI feeds these files directly to the model as prompts, producing deterministic outputs that can be asserted against like traditional API responses. This eliminates non-deterministic behavior by enforcing hard rules that abort execution when violated.

Can I use any testing framework with the skills CLI?

Yes. Because the library exposes functionality through a standard CLI (npx skills@latest), you can invoke it from Jest, Pytest, Go tests, or shell scripts. Simply capture stdout from commands like npx skills@latest run animate and assert against the output strings or JSON structures.

How do I test gate rejection scenarios deliberately?

Craft a request that violates the hard rules defined in the skill's markdown. For example, when testing skills/animate/SKILL.md, submit the request "animate a keyboard shortcut" if that action is prohibited by the hard rules. Assert that the CLI output contains a gate-rejection message rather than generated code, confirming that the automatic safety mechanisms function correctly.

Where are the skill definitions stored in the repository?

Skill definitions reside in the skills/ directory as markdown files. Each skill has its own subdirectory containing a SKILL.md file (such as skills/pick-ui-library/SKILL.md or skills/improve-animations/SKILL.md) that defines the metadata, workflow steps, and hard rules. Auxiliary resources like skills/animate/RECIPES.md provide additional reference material that the agent fetches at runtime.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →