# How the Shared Instruction Builder in Ponytail Works: Mode-Aware LLM Prompting

> Discover how Ponytail's shared instruction builder enhances LLM prompting. Learn how it injects mode-specific guidance for lite, full, or ultra runtime modes.

- Repository: [DietrichGebert/ponytail](https://github.com/DietrichGebert/ponytail)
- Tags: deep-dive
- Published: 2026-09-13

---

**Ponytail's shared instruction builder dynamically injects mode-specific guidance into every LLM turn by loading [`skills/ponytail/SKILL.md`](https://github.com/DietrichGebert/ponytail/blob/main/skills/ponytail/SKILL.md), stripping YAML front-matter, filtering content based on the active runtime mode (lite, full, ultra, or off), and returning formatted instructions via `build_injected_context()`.**

The DietrichGebert/ponytail repository implements a deterministic approach to agent consistency through its shared instruction builder—a self-contained pipeline defined in [`__init__.py`](https://github.com/DietrichGebert/ponytail/blob/main/__init__.py) that ensures every LLM interaction receives appropriately scoped guidance based on the current operational mode.

## Core Architecture of the Instruction Pipeline

The shared instruction builder operates as a single source of truth for all agent prompting, transforming static markdown files into runtime-specific context blocks. The pipeline executes seven distinct steps to construct the final instruction string injected before each LLM call.

### Runtime Mode Resolution

The system normalizes the effective mode from three potential sources: environment variables, user-supplied arguments, or the default configuration. In [`__init__.py`](https://github.com/DietrichGebert/ponytail/blob/main/__init__.py) (lines 30–63), the functions `_normalize_runtime_mode` and `_normalize_config_mode` validate inputs against supported modes—**lite**, **full**, **ultra**, and **off**—falling back to `_default_mode` when unspecified.

### Instruction Source Loading

The human-readable guidance resides in [`skills/ponytail/SKILL.md`](https://github.com/DietrichGebert/ponytail/blob/main/skills/ponytail/SKILL.md). The `build_injected_context` function (lines 17–21) calls `PONYTAIL_SKILL.read_text()` to load the raw markdown content into memory. For review-specific operations, the system alternatively sources from [`skills/ponytail-review/SKILL.md`](https://github.com/DietrichGebert/ponytail/blob/main/skills/ponytail-review/SKILL.md).

### Content Processing Pipeline

Raw markdown undergoes two transformation stages before injection:

- **Front-matter stripping**: The `_strip_frontmatter` function (lines 66–68) removes YAML metadata blocks from the top of the file, ensuring only instructional content reaches the LLM.

- **Mode-specific filtering**: The `_filter_skill_body_for_mode` function (lines 70–87) parses each line for mode prefixes (e.g., `| **lite** |` or `- lite:`). It discards sections incompatible with the effective mode, returning only relevant guidance for the current operational context.

### Fallback Mechanisms

If the skill file cannot be read or is missing, the builder activates a built-in fallback defined in lines 90–102. This fallback encodes Ponytail's core "lazy senior developer" philosophy, ensuring agents retain baseline behavioral guidelines even when external files are unavailable.

## LLM Integration and Context Injection

The `_pre_llm_call` function (lines 125–129) serves as the Hermès integration hook. It obtains the current mode, triggers `build_injected_context()`, and returns a dictionary structure: `{"context": …}`. This dict is consumed by the LLM runtime, placing the filtered instruction block at the beginning of every turn to establish consistent agent behavior.

## Implementing Mode-Aware Instructions

The following examples demonstrate practical usage of the shared instruction builder API:

```python
from ponytail import build_injected_context, _pre_llm_call

# Retrieve instructions for the default mode (full)

default_context = build_injected_context()
print(default_context)

# Output: "PONYTAIL MODE ACTIVE — level: full\n\n…filtered skill body…"

# Force lite mode for constrained context windows

lite_context = build_injected_context("lite")
print(lite_context)

# Output: "PONYTAIL MODE ACTIVE — level: lite\n\n…only lite‑specific sections…"

# Simulate the Hermès pre-LLM hook

hook_result = _pre_llm_call(session_id="123")
print(hook_result["context"])

# Returns the same mode-filtered string as build_injected_context()

```

## Summary

- The shared instruction builder lives in [`__init__.py`](https://github.com/DietrichGebert/ponytail/blob/main/__init__.py) and processes [`skills/ponytail/SKILL.md`](https://github.com/DietrichGebert/ponytail/blob/main/skills/ponytail/SKILL.md) to generate consistent agent guidance.
- Mode resolution supports **lite**, **full**, **ultra**, and **off** settings via `_normalize_runtime_mode` and `_default_mode`.
- Content filtering removes YAML front-matter and extracts only mode-relevant sections using `_filter_skill_body_for_mode`.
- A built-in fallback ensures operational continuity when skill files are missing.
- The `_pre_llm_call` hook integrates with the Hermès framework to inject context before every LLM turn.

## Frequently Asked Questions

### What file contains the raw instruction templates for Ponytail?

The primary instruction source is [`skills/ponytail/SKILL.md`](https://github.com/DietrichGebert/ponytail/blob/main/skills/ponytail/SKILL.md) in the DietrichGebert/ponytail repository. This markdown file contains the human-written guidance that the shared instruction builder loads, filters, and injects into LLM contexts.

### How does Ponytail handle invalid or missing mode configurations?

When the specified mode cannot be resolved, the system falls back to `_default_mode` as defined in [`__init__.py`](https://github.com/DietrichGebert/ponytail/blob/main/__init__.py) (lines 30–63). If the skill file itself is unreadable, the builder uses a hardcoded fallback description of the "lazy senior developer" philosophy (lines 90–102) to maintain baseline agent behavior.

### How does the instruction builder integrate with the Hermès agent framework?

The `_pre_llm_call` function (lines 125–129) serves as the integration point. It constructs the mode-filtered context and returns it as a dictionary with a `"context"` key, which the Hermès runtime consumes to prepend instructions to each LLM request.

### Can the instruction content be customized for specific operational modes?

Yes. The `_filter_skill_body_for_mode` function (lines 70–87) specifically searches for mode-labeled sections within [`SKILL.md`](https://github.com/DietrichGebert/ponytail/blob/main/SKILL.md) (such as `| **lite** |` headers). Content creators can add mode-specific blocks to the markdown file, and the builder will automatically include only the sections matching the active runtime mode.