# How Hyperresearch Persists and Uses the Canonical Research Query

> Discover how Hyperresearch stores and utilizes the canonical research query in query.md files. Learn how agents, linters, and hooks reference the exact source question for consistent pipeline execution.

- Repository: [Jordan Gibbs/hyperresearch](https://github.com/jordan-gibbs/hyperresearch)
- Tags: internals
- Published: 2026-09-13

---

**In Hyperresearch, the canonical research query is stored as a [`query.md`](https://github.com/jordan-gibbs/hyperresearch/blob/main/query.md) file within each run's workspace at `research/runs/<vault_tag>/query.md`, containing the verbatim user prompt with YAML front-matter, and is read by agents, linters, and hooks to ensure every pipeline stage references the exact same source question.**

Hyperresearch is an open-source research automation framework designed to ensure reproducible, traceable AI-assisted investigation. At the heart of every research run lies the **canonical research query**—the immutable record of exactly what the user asked the system to investigate. This article examines how the jordan-gibbs/hyperresearch repository persists this query to disk and consumes it across the entire pipeline.

## Persistence Layer: Where the Canonical Query Lives

The system writes the canonical query to a predictable filesystem location immediately upon run initialization. For standard runs, the file resides at:

```

research/runs/<vault_tag>/query.md

```

This path follows a strict convention where `<vault_tag>` represents the unique identifier for the research run. The [`query.md`](https://github.com/jordan-gibbs/hyperresearch/blob/main/query.md) file contains the verbatim user prompt prepended with a small YAML front-matter block capturing metadata such as timestamps and tags.

For **wrapped runs**—executions orchestrated by external harnesses—the framework checks for [`research/prompt.txt`](https://github.com/jordan-gibbs/hyperresearch/blob/main/research/prompt.txt) at the repository root. If present, this file takes precedence and overwrites any other prompt source, serving as the gospel authority for that run's canonical query.

## Run Initialization: Creating the Canonical Query

The scaffolding logic that establishes the per-run workspace and writes the initial query file lives in **[`hyperresearch/core/runs.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/hyperresearch/core/runs.py)**. When the `RunScaffold` routine creates a new run directory, it performs the following:

1. Creates the workspace structure under `research/runs/<vault_tag>/`
2. Generates the [`query.md`](https://github.com/jordan-gibbs/hyperresearch/blob/main/query.md) file with YAML front-matter
3. Persists the user-supplied prompt text into this file

This single write operation establishes the immutable source of truth that all downstream components will reference.

## Consumption Patterns: How the Pipeline Reads the Query

Once persisted, the canonical research query becomes the authoritative reference for every subsequent pipeline stage. The framework employs a consistent access pattern—utilizing `Vault.run_dir(tag) / "query.md"`—to ensure all components read the identical text.

### Agent Documentation Injection

In **[`hyperresearch/core/agent_docs.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/hyperresearch/core/agent_docs.py)**, the system reads `research/runs/<vault_tag>/query.md` to obtain the canonical query. This content is injected directly into LLM prompts for all downstream agents, ensuring that every AI-assisted step operates with full context of the original research question.

### Linting and Validation

The **[`hyperresearch/cli/lint.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/hyperresearch/cli/lint.py)** module performs validation by loading the canonical prompt from either [`query.md`](https://github.com/jordan-gibbs/hyperresearch/blob/main/query.md) or [`research/prompt.txt`](https://github.com/jordan-gibbs/hyperresearch/blob/main/research/prompt.txt). It compares this stored value against the prompt extracted from the scaffolded run. If the lint rule `wrapper-report` detects a mismatch between the stored `canonical_prompt` and the extracted version, it raises an assertion error, preventing silent prompt corruption.

### Report Generation and Hooks

Various hooks in **[`hyperresearch/core/hooks.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/hyperresearch/core/hooks.py)** reference the canonical query when building final reports, checking for contradictions, or generating scaffold sections. Throughout the entire lifecycle, code consistently accesses the query via the vault helper to guarantee that every artifact traces back to the exact original question.

## Practical Implementation Examples

```python

# Example: Load the canonical query for a given vault tag

from pathlib import Path

def load_canonical_query(vault, tag: str) -> str:
    """Return the verbatim research query for the run identified by `tag`."""
    query_path = vault.run_dir(tag) / "query.md"
    if not query_path.is_file():
        raise FileNotFoundError(f"No canonical query found at {query_path}")
    # The file starts with YAML front-matter; split it off.

    content = query_path.read_text(encoding="utf-8")
    _, prompt = content.split("---", maxsplit=2)[1:]  # strip front-matter

    return prompt.strip()

```

```python

# Example: Lint rule that verifies the stored canonical prompt matches the scaffolded one

def _check_canonical_prompt(vault, tag, extracted_prompt):
    canonical_path = vault.run_dir(tag) / "query.md"
    if canonical_path.is_file():
        canonical_prompt = canonical_path.read_text().split("---", 2)[2].strip()
        if extracted_prompt != canonical_prompt:
            raise AssertionError(
                f"Canonical prompt mismatch for tag {tag}: "
                f"expected '{canonical_prompt}', got '{extracted_prompt}'"
            )

```

## Key Files and Responsibilities

The canonical query system spans several core modules:

- **[`src/hyperresearch/core/runs.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/runs.py)**: Creates the per-run workspace and writes [`query.md`](https://github.com/jordan-gibbs/hyperresearch/blob/main/query.md) during scaffold initialization.
- **[`src/hyperresearch/core/agent_docs.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/agent_docs.py)**: Reads [`query.md`](https://github.com/jordan-gibbs/hyperresearch/blob/main/query.md) to inject the canonical query into LLM prompts for downstream agents.
- **[`src/hyperresearch/cli/lint.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/cli/lint.py)**: Validates that the stored canonical prompt matches the scaffolded prompt via the `wrapper-report` rule.
- **[`src/hyperresearch/core/vault.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/vault.py)**: Provides the `Vault.run_dir(tag)` helper used throughout the codebase to locate `research/runs/<vault_tag>/`.
- **[`src/hyperresearch/core/hooks.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/hooks.py)**: References the canonical query during report generation and contradiction checking.

## Summary

- The **canonical research query** is persisted to `research/runs/<vault_tag>/query.md` immediately upon run initialization, or to [`research/prompt.txt`](https://github.com/jordan-gibbs/hyperresearch/blob/main/research/prompt.txt) for wrapped runs.
- **[`hyperresearch/core/runs.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/hyperresearch/core/runs.py)** handles the initial scaffolding and write operations via the `RunScaffold` routine.
- **[`hyperresearch/core/agent_docs.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/hyperresearch/core/agent_docs.py)**, **[`hyperresearch/cli/lint.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/hyperresearch/cli/lint.py)**, and **[`hyperresearch/core/hooks.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/hyperresearch/core/hooks.py)** all consume this file to ensure consistent reference to the original user prompt.
- The system guarantees reproducibility by using `Vault.run_dir(tag) / "query.md"` as the single source of truth across all pipeline stages.

## Frequently Asked Questions

### What is the canonical research query in Hyperresearch?

The canonical research query is the immutable, verbatim record of the user's original prompt that initiated a research run. It serves as the single source of truth for what the system was asked to investigate, persisted to disk to ensure all pipeline stages reference the exact same question.

### How does Hyperresearch handle prompt changes during wrapped runs?

During wrapped runs, Hyperresearch checks for [`research/prompt.txt`](https://github.com/jordan-gibbs/hyperresearch/blob/main/research/prompt.txt) at the repository root. If this file exists, it takes precedence over the standard [`query.md`](https://github.com/jordan-gibbs/hyperresearch/blob/main/query.md) location and overwrites any other prompt source, ensuring external harnesses cannot silently modify the canonical query without updating the persisted file.

### Where is the canonical query stored in the filesystem?

For standard runs, the canonical query resides at `research/runs/<vault_tag>/query.md` within the run-specific workspace. For wrapped runs, it may alternatively be stored at [`research/prompt.txt`](https://github.com/jordan-gibbs/hyperresearch/blob/main/research/prompt.txt). The `Vault.run_dir(tag)` helper in **[`hyperresearch/core/vault.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/hyperresearch/core/vault.py)** provides the canonical path resolution.

### How does the linting system verify the canonical query?

The **[`hyperresearch/cli/lint.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/hyperresearch/cli/lint.py)** module loads the canonical prompt from [`query.md`](https://github.com/jordan-gibbs/hyperresearch/blob/main/query.md) or [`research/prompt.txt`](https://github.com/jordan-gibbs/hyperresearch/blob/main/research/prompt.txt) and compares it against the prompt extracted from the scaffolded run. The `wrapper-report` lint rule raises an assertion error if a mismatch is detected, preventing inconsistencies between the stored query and the active run configuration.