How Hyperresearch Handles Open-Access Recovery for Paywalled Papers
Hyperresearch automates open-access recovery for paywalled papers by querying Unpaywall, Europe PMC, and CORE to rescue or substitute full-text content when the original source is inaccessible or provides only thin abstracts.
The jordan-gibbs/hyperresearch repository implements a robust open-access recovery system that retrieves legal full-text copies when academic papers are behind paywalls. By integrating three major open-access APIs, it ensures researchers can access content without violating copyright while maintaining complete transparency about the provenance of recovered materials.
The Two-Stage Open-Access Recovery Pipeline
Hyperresearch distinguishes between completely blocked content and partially accessible papers through two distinct recovery modes implemented in src/hyperresearch/core/oa.py.
Stage 1: Rescue for Inaccessible URLs
When a fetch operation fails due to paywalls, authentication requirements, or bot detection (HTTP 403/401 errors), Hyperresearch invokes rescue_full_text (lines 64‑85). This function extracts the DOI from the original URL and iterates through open-access candidates until it finds a readable full-text. The resulting content is stored as a rescued note, meaning the original source was never read or processed.
Stage 2: Substitute for Thin Abstracts
If the initial fetch returns a landing page or thin abstract instead of full text, recover_full_text (lines 40‑49) evaluates whether the result needs_oa_recovery. When triggered, this function queries the same OA APIs but performs a length and quality comparison. It only substitutes the original content if the open-access version is substantially longer and passes validation. These notes are marked as substituted, preserving the original URL as the source while appending a disclosure banner.
How OA Candidates Are Discovered and Validated
Before any network request, Hyperresearch validates potential sources through a tiered discovery and security pipeline.
The Candidate Iterator
The iter_oa_candidates function (lines 22‑34) implements a lazy generator that queries Unpaywall, Europe PMC, and CORE in order of preference. It yields candidates prioritized by format: PDFs first, followed by landing pages, and finally CORE plain-text records. This ensures the highest-fidelity copy is retrieved first without unnecessary API calls.
Security and Safety Checks
Every candidate URL undergoes validation by check_oa_url (lines 34‑50) to prevent Server-Side Request Forgery (SSRF) and private network access. The function parses the URL, confirms an HTTP(S) scheme, and verifies that the resolved IP address is public before allowing any fetch operation to proceed.
The Three Open-Access Resolvers
Hyperresearch aggregates open-access content through three specialized providers, each handling different domains and authentication requirements.
Unpaywall Integration
Unpaywall provides broad coverage across all academic disciplines and requires a contact email configured in settings.contact_email. The _unpaywall_candidates function (lines 97‑124) parses the API response to build a sorted list of locations, prioritizing PDF links and ranking versions by quality (publishedVersion preferred over acceptedVersion).
Europe PMC for Biomedical Literature
For biomedical and life sciences papers, _resolve_europepmc (lines 223‑252) queries Europe PMC, which requires no API key. This resolver specifically handles JATS XML (fullTextXML) responses, converting structured XML into markdown format. It returns an OALocation with kind jats when successful.
CORE Aggregator
CORE aggregates the largest collection of open-access records and requires an API key via the CORE_API_KEY environment variable. The _core_candidates function (lines 404‑447) yields copies in priority order: plain-text (coretext) first, then hosted PDFs, and finally harvested URLs from the CORE repository.
CLI Integration and Automatic Recovery
The recovery pipeline integrates seamlessly into the standard fetch workflow through the CLI interface.
When executing hyperresearch fetch or hyperresearch fetch-batch, the system automatically attempts recovery as implemented in src/hyperresearch/cli/fetch.py (lines 13‑26). If a fetch raises an exception, the CLI invokes the rescue pathway. If the fetch succeeds but returns thin content, it triggers the substitute pathway.
# Single paper fetch with automatic OA recovery
hyperresearch fetch https://doi.org/10.1038/s41586-020-2649-2
For custom scripts, the Python API exposes both recovery modes directly:
from hyperresearch.core.oa import rescue_full_text, recover_full_text
# Initialize vault and provider
vault = ... # hyperresearch.Vault instance
provider = ... # Web provider (e.g., TavilyProvider)
url = "https://example.com/paywalled-paper"
doi = "10.1234/example.doi"
# Rescue mode: source never read
rescued, location = rescue_full_text(vault, provider, url, doi)
if rescued:
print(f"Rescued from: {location.url}")
# Substitute mode: replace thin abstract
thin_result = provider.fetch(url)
full_text, location = recover_full_text(
vault, provider, url, doi, thin_result
)
if full_text:
print(f"Substituted with: {location.url}")
Metadata Transparency and Disclosure
When Hyperresearch uses an open-access copy, it maintains strict provenance tracking through banner generation and front-matter metadata.
The recovery_notice function (lines 66‑82) generates a markdown blockquote banner that appears at the top of the note, disclosing that the content was recovered from an alternative source. Simultaneously, oa_frontmatter injects structured YAML fields (oa_source, oa_resolver, oa_version, oa_license) into the note's metadata, ensuring complete auditability of the recovery chain.
Summary
- Hyperresearch implements two-stage open-access recovery: rescue for blocked URLs and substitute for thin abstracts.
- The system queries Unpaywall, Europe PMC, and CORE through
iter_oa_candidatesinsrc/hyperresearch/core/oa.py. - Security validation via
check_oa_urlprevents SSRF by verifying public IP ranges before fetching. - Recovery triggers automatically in
src/hyperresearch/cli/fetch.pyduring standard fetch operations. - All recovered content includes provenance metadata and disclosure banners to maintain transparency.
Frequently Asked Questions
What triggers open-access recovery in Hyperresearch?
Recovery triggers in two scenarios: when rescue_full_text detects a fetch exception (HTTP 403/401 or connection errors) in src/hyperresearch/cli/fetch.py, or when recover_full_text identifies content as "thin" (shorter than expected full-text) after a successful fetch. Both paths query the three OA APIs to find legal alternatives.
Which open-access databases does Hyperresearch search?
The system searches Unpaywall (general coverage, requires contact email), Europe PMC (biomedical focus, no key required), and CORE (largest aggregator, requires API key). The iter_oa_candidates function queries them sequentially in that order.
How does Hyperresearch prevent security risks when fetching OA copies?
Before downloading any candidate, check_oa_url (lines 34‑50 in src/hyperresearch/core/oa.py) parses the URL, enforces HTTP(S) protocols, and resolves the IP address to confirm it is public. This blocks private network ranges and prevents SSRF attacks against internal infrastructure.
What is the difference between rescued and substituted notes?
Rescued notes occur when the original URL is completely inaccessible; the system reads nothing from the source and replaces it entirely with OA content. Substituted notes occur when the original provides a thin abstract; the system preserves the original URL as the source but appends a banner and metadata indicating the full text came from an open-access repository.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →