How Hyperresearch Persists and Uses the Canonical Research Query

In Hyperresearch, the canonical research query is stored as a query.md file within each run's workspace at research/runs/<vault_tag>/query.md, containing the verbatim user prompt with YAML front-matter, and is read by agents, linters, and hooks to ensure every pipeline stage references the exact same source question.

Hyperresearch is an open-source research automation framework designed to ensure reproducible, traceable AI-assisted investigation. At the heart of every research run lies the canonical research query—the immutable record of exactly what the user asked the system to investigate. This article examines how the jordan-gibbs/hyperresearch repository persists this query to disk and consumes it across the entire pipeline.

Persistence Layer: Where the Canonical Query Lives

The system writes the canonical query to a predictable filesystem location immediately upon run initialization. For standard runs, the file resides at:


research/runs/<vault_tag>/query.md

This path follows a strict convention where <vault_tag> represents the unique identifier for the research run. The query.md file contains the verbatim user prompt prepended with a small YAML front-matter block capturing metadata such as timestamps and tags.

For wrapped runs—executions orchestrated by external harnesses—the framework checks for research/prompt.txt at the repository root. If present, this file takes precedence and overwrites any other prompt source, serving as the gospel authority for that run's canonical query.

Run Initialization: Creating the Canonical Query

The scaffolding logic that establishes the per-run workspace and writes the initial query file lives in hyperresearch/core/runs.py. When the RunScaffold routine creates a new run directory, it performs the following:

  1. Creates the workspace structure under research/runs/<vault_tag>/
  2. Generates the query.md file with YAML front-matter
  3. Persists the user-supplied prompt text into this file

This single write operation establishes the immutable source of truth that all downstream components will reference.

Consumption Patterns: How the Pipeline Reads the Query

Once persisted, the canonical research query becomes the authoritative reference for every subsequent pipeline stage. The framework employs a consistent access pattern—utilizing Vault.run_dir(tag) / "query.md"—to ensure all components read the identical text.

Agent Documentation Injection

In hyperresearch/core/agent_docs.py, the system reads research/runs/<vault_tag>/query.md to obtain the canonical query. This content is injected directly into LLM prompts for all downstream agents, ensuring that every AI-assisted step operates with full context of the original research question.

Linting and Validation

The hyperresearch/cli/lint.py module performs validation by loading the canonical prompt from either query.md or research/prompt.txt. It compares this stored value against the prompt extracted from the scaffolded run. If the lint rule wrapper-report detects a mismatch between the stored canonical_prompt and the extracted version, it raises an assertion error, preventing silent prompt corruption.

Report Generation and Hooks

Various hooks in hyperresearch/core/hooks.py reference the canonical query when building final reports, checking for contradictions, or generating scaffold sections. Throughout the entire lifecycle, code consistently accesses the query via the vault helper to guarantee that every artifact traces back to the exact original question.

Practical Implementation Examples


# Example: Load the canonical query for a given vault tag

from pathlib import Path

def load_canonical_query(vault, tag: str) -> str:
    """Return the verbatim research query for the run identified by `tag`."""
    query_path = vault.run_dir(tag) / "query.md"
    if not query_path.is_file():
        raise FileNotFoundError(f"No canonical query found at {query_path}")
    # The file starts with YAML front-matter; split it off.

    content = query_path.read_text(encoding="utf-8")
    _, prompt = content.split("---", maxsplit=2)[1:]  # strip front-matter

    return prompt.strip()

# Example: Lint rule that verifies the stored canonical prompt matches the scaffolded one

def _check_canonical_prompt(vault, tag, extracted_prompt):
    canonical_path = vault.run_dir(tag) / "query.md"
    if canonical_path.is_file():
        canonical_prompt = canonical_path.read_text().split("---", 2)[2].strip()
        if extracted_prompt != canonical_prompt:
            raise AssertionError(
                f"Canonical prompt mismatch for tag {tag}: "
                f"expected '{canonical_prompt}', got '{extracted_prompt}'"
            )

Key Files and Responsibilities

The canonical query system spans several core modules:

Summary

Frequently Asked Questions

What is the canonical research query in Hyperresearch?

The canonical research query is the immutable, verbatim record of the user's original prompt that initiated a research run. It serves as the single source of truth for what the system was asked to investigate, persisted to disk to ensure all pipeline stages reference the exact same question.

How does Hyperresearch handle prompt changes during wrapped runs?

During wrapped runs, Hyperresearch checks for research/prompt.txt at the repository root. If this file exists, it takes precedence over the standard query.md location and overwrites any other prompt source, ensuring external harnesses cannot silently modify the canonical query without updating the persisted file.

Where is the canonical query stored in the filesystem?

For standard runs, the canonical query resides at research/runs/<vault_tag>/query.md within the run-specific workspace. For wrapped runs, it may alternatively be stored at research/prompt.txt. The Vault.run_dir(tag) helper in hyperresearch/core/vault.py provides the canonical path resolution.

How does the linting system verify the canonical query?

The hyperresearch/cli/lint.py module loads the canonical prompt from query.md or research/prompt.txt and compares it against the prompt extracted from the scaffolded run. The wrapper-report lint rule raises an assertion error if a mismatch is detected, preventing inconsistencies between the stored query and the active run configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →