How Hyperresearch Persists and Uses the Canonical Research Query
In Hyperresearch, the canonical research query is stored as a query.md file within each run's workspace at research/runs/<vault_tag>/query.md, containing the verbatim user prompt with YAML front-matter, and is read by agents, linters, and hooks to ensure every pipeline stage references the exact same source question.
Hyperresearch is an open-source research automation framework designed to ensure reproducible, traceable AI-assisted investigation. At the heart of every research run lies the canonical research query—the immutable record of exactly what the user asked the system to investigate. This article examines how the jordan-gibbs/hyperresearch repository persists this query to disk and consumes it across the entire pipeline.
Persistence Layer: Where the Canonical Query Lives
The system writes the canonical query to a predictable filesystem location immediately upon run initialization. For standard runs, the file resides at:
research/runs/<vault_tag>/query.md
This path follows a strict convention where <vault_tag> represents the unique identifier for the research run. The query.md file contains the verbatim user prompt prepended with a small YAML front-matter block capturing metadata such as timestamps and tags.
For wrapped runs—executions orchestrated by external harnesses—the framework checks for research/prompt.txt at the repository root. If present, this file takes precedence and overwrites any other prompt source, serving as the gospel authority for that run's canonical query.
Run Initialization: Creating the Canonical Query
The scaffolding logic that establishes the per-run workspace and writes the initial query file lives in hyperresearch/core/runs.py. When the RunScaffold routine creates a new run directory, it performs the following:
- Creates the workspace structure under
research/runs/<vault_tag>/ - Generates the
query.mdfile with YAML front-matter - Persists the user-supplied prompt text into this file
This single write operation establishes the immutable source of truth that all downstream components will reference.
Consumption Patterns: How the Pipeline Reads the Query
Once persisted, the canonical research query becomes the authoritative reference for every subsequent pipeline stage. The framework employs a consistent access pattern—utilizing Vault.run_dir(tag) / "query.md"—to ensure all components read the identical text.
Agent Documentation Injection
In hyperresearch/core/agent_docs.py, the system reads research/runs/<vault_tag>/query.md to obtain the canonical query. This content is injected directly into LLM prompts for all downstream agents, ensuring that every AI-assisted step operates with full context of the original research question.
Linting and Validation
The hyperresearch/cli/lint.py module performs validation by loading the canonical prompt from either query.md or research/prompt.txt. It compares this stored value against the prompt extracted from the scaffolded run. If the lint rule wrapper-report detects a mismatch between the stored canonical_prompt and the extracted version, it raises an assertion error, preventing silent prompt corruption.
Report Generation and Hooks
Various hooks in hyperresearch/core/hooks.py reference the canonical query when building final reports, checking for contradictions, or generating scaffold sections. Throughout the entire lifecycle, code consistently accesses the query via the vault helper to guarantee that every artifact traces back to the exact original question.
Practical Implementation Examples
# Example: Load the canonical query for a given vault tag
from pathlib import Path
def load_canonical_query(vault, tag: str) -> str:
"""Return the verbatim research query for the run identified by `tag`."""
query_path = vault.run_dir(tag) / "query.md"
if not query_path.is_file():
raise FileNotFoundError(f"No canonical query found at {query_path}")
# The file starts with YAML front-matter; split it off.
content = query_path.read_text(encoding="utf-8")
_, prompt = content.split("---", maxsplit=2)[1:] # strip front-matter
return prompt.strip()
# Example: Lint rule that verifies the stored canonical prompt matches the scaffolded one
def _check_canonical_prompt(vault, tag, extracted_prompt):
canonical_path = vault.run_dir(tag) / "query.md"
if canonical_path.is_file():
canonical_prompt = canonical_path.read_text().split("---", 2)[2].strip()
if extracted_prompt != canonical_prompt:
raise AssertionError(
f"Canonical prompt mismatch for tag {tag}: "
f"expected '{canonical_prompt}', got '{extracted_prompt}'"
)
Key Files and Responsibilities
The canonical query system spans several core modules:
src/hyperresearch/core/runs.py: Creates the per-run workspace and writesquery.mdduring scaffold initialization.src/hyperresearch/core/agent_docs.py: Readsquery.mdto inject the canonical query into LLM prompts for downstream agents.src/hyperresearch/cli/lint.py: Validates that the stored canonical prompt matches the scaffolded prompt via thewrapper-reportrule.src/hyperresearch/core/vault.py: Provides theVault.run_dir(tag)helper used throughout the codebase to locateresearch/runs/<vault_tag>/.src/hyperresearch/core/hooks.py: References the canonical query during report generation and contradiction checking.
Summary
- The canonical research query is persisted to
research/runs/<vault_tag>/query.mdimmediately upon run initialization, or toresearch/prompt.txtfor wrapped runs. hyperresearch/core/runs.pyhandles the initial scaffolding and write operations via theRunScaffoldroutine.hyperresearch/core/agent_docs.py,hyperresearch/cli/lint.py, andhyperresearch/core/hooks.pyall consume this file to ensure consistent reference to the original user prompt.- The system guarantees reproducibility by using
Vault.run_dir(tag) / "query.md"as the single source of truth across all pipeline stages.
Frequently Asked Questions
What is the canonical research query in Hyperresearch?
The canonical research query is the immutable, verbatim record of the user's original prompt that initiated a research run. It serves as the single source of truth for what the system was asked to investigate, persisted to disk to ensure all pipeline stages reference the exact same question.
How does Hyperresearch handle prompt changes during wrapped runs?
During wrapped runs, Hyperresearch checks for research/prompt.txt at the repository root. If this file exists, it takes precedence over the standard query.md location and overwrites any other prompt source, ensuring external harnesses cannot silently modify the canonical query without updating the persisted file.
Where is the canonical query stored in the filesystem?
For standard runs, the canonical query resides at research/runs/<vault_tag>/query.md within the run-specific workspace. For wrapped runs, it may alternatively be stored at research/prompt.txt. The Vault.run_dir(tag) helper in hyperresearch/core/vault.py provides the canonical path resolution.
How does the linting system verify the canonical query?
The hyperresearch/cli/lint.py module loads the canonical prompt from query.md or research/prompt.txt and compares it against the prompt extracted from the scaffolded run. The wrapper-report lint rule raises an assertion error if a mismatch is detected, preventing inconsistencies between the stored query and the active run configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →