# How Hyperresearch Handles Open-Access Recovery for Paywalled Papers

> Discover how Hyperresearch automatically recovers open-access versions of paywalled papers. Learn how it queries Unpaywall, Europe PMC, and CORE to find full-text content.

- Repository: [Jordan Gibbs/hyperresearch](https://github.com/jordan-gibbs/hyperresearch)
- Tags: how-to-guide
- Published: 2026-09-13

---

**Hyperresearch automates open-access recovery for paywalled papers by querying Unpaywall, Europe PMC, and CORE to rescue or substitute full-text content when the original source is inaccessible or provides only thin abstracts.**

The `jordan-gibbs/hyperresearch` repository implements a robust open-access recovery system that retrieves legal full-text copies when academic papers are behind paywalls. By integrating three major open-access APIs, it ensures researchers can access content without violating copyright while maintaining complete transparency about the provenance of recovered materials.

## The Two-Stage Open-Access Recovery Pipeline

Hyperresearch distinguishes between completely blocked content and partially accessible papers through two distinct recovery modes implemented in [`src/hyperresearch/core/oa.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/oa.py).

### Stage 1: Rescue for Inaccessible URLs

When a fetch operation fails due to paywalls, authentication requirements, or bot detection (HTTP 403/401 errors), Hyperresearch invokes `rescue_full_text` (lines 64‑85). This function extracts the DOI from the original URL and iterates through open-access candidates until it finds a readable full-text. The resulting content is stored as a **rescued** note, meaning the original source was never read or processed.

### Stage 2: Substitute for Thin Abstracts

If the initial fetch returns a landing page or thin abstract instead of full text, `recover_full_text` (lines 40‑49) evaluates whether the result `needs_oa_recovery`. When triggered, this function queries the same OA APIs but performs a length and quality comparison. It only substitutes the original content if the open-access version is substantially longer and passes validation. These notes are marked as **substituted**, preserving the original URL as the source while appending a disclosure banner.

## How OA Candidates Are Discovered and Validated

Before any network request, Hyperresearch validates potential sources through a tiered discovery and security pipeline.

### The Candidate Iterator

The `iter_oa_candidates` function (lines 22‑34) implements a lazy generator that queries Unpaywall, Europe PMC, and CORE in order of preference. It yields candidates prioritized by format: PDFs first, followed by landing pages, and finally CORE plain-text records. This ensures the highest-fidelity copy is retrieved first without unnecessary API calls.

### Security and Safety Checks

Every candidate URL undergoes validation by `check_oa_url` (lines 34‑50) to prevent Server-Side Request Forgery (SSRF) and private network access. The function parses the URL, confirms an HTTP(S) scheme, and verifies that the resolved IP address is public before allowing any fetch operation to proceed.

## The Three Open-Access Resolvers

Hyperresearch aggregates open-access content through three specialized providers, each handling different domains and authentication requirements.

### Unpaywall Integration

Unpaywall provides broad coverage across all academic disciplines and requires a contact email configured in `settings.contact_email`. The `_unpaywall_candidates` function (lines 97‑124) parses the API response to build a sorted list of locations, prioritizing PDF links and ranking versions by quality (`publishedVersion` preferred over `acceptedVersion`).

### Europe PMC for Biomedical Literature

For biomedical and life sciences papers, `_resolve_europepmc` (lines 223‑252) queries Europe PMC, which requires no API key. This resolver specifically handles JATS XML (`fullTextXML`) responses, converting structured XML into markdown format. It returns an `OALocation` with kind `jats` when successful.

### CORE Aggregator

CORE aggregates the largest collection of open-access records and requires an API key via the `CORE_API_KEY` environment variable. The `_core_candidates` function (lines 404‑447) yields copies in priority order: plain-text (`coretext`) first, then hosted PDFs, and finally harvested URLs from the CORE repository.

## CLI Integration and Automatic Recovery

The recovery pipeline integrates seamlessly into the standard fetch workflow through the CLI interface.

When executing `hyperresearch fetch` or `hyperresearch fetch-batch`, the system automatically attempts recovery as implemented in [`src/hyperresearch/cli/fetch.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/cli/fetch.py) (lines 13‑26). If a fetch raises an exception, the CLI invokes the rescue pathway. If the fetch succeeds but returns thin content, it triggers the substitute pathway.

```bash

# Single paper fetch with automatic OA recovery

hyperresearch fetch https://doi.org/10.1038/s41586-020-2649-2

```

For custom scripts, the Python API exposes both recovery modes directly:

```python
from hyperresearch.core.oa import rescue_full_text, recover_full_text

# Initialize vault and provider

vault = ...          # hyperresearch.Vault instance

provider = ...       # Web provider (e.g., TavilyProvider)

url = "https://example.com/paywalled-paper"
doi = "10.1234/example.doi"

# Rescue mode: source never read

rescued, location = rescue_full_text(vault, provider, url, doi)
if rescued:
    print(f"Rescued from: {location.url}")

# Substitute mode: replace thin abstract

thin_result = provider.fetch(url)
full_text, location = recover_full_text(
    vault, provider, url, doi, thin_result
)
if full_text:
    print(f"Substituted with: {location.url}")

```

## Metadata Transparency and Disclosure

When Hyperresearch uses an open-access copy, it maintains strict provenance tracking through banner generation and front-matter metadata.

The `recovery_notice` function (lines 66‑82) generates a markdown blockquote banner that appears at the top of the note, disclosing that the content was recovered from an alternative source. Simultaneously, `oa_frontmatter` injects structured YAML fields (`oa_source`, `oa_resolver`, `oa_version`, `oa_license`) into the note's metadata, ensuring complete auditability of the recovery chain.

## Summary

- Hyperresearch implements **two-stage open-access recovery**: rescue for blocked URLs and substitute for thin abstracts.
- The system queries **Unpaywall, Europe PMC, and CORE** through `iter_oa_candidates` in [`src/hyperresearch/core/oa.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/oa.py).
- **Security validation** via `check_oa_url` prevents SSRF by verifying public IP ranges before fetching.
- Recovery triggers automatically in [`src/hyperresearch/cli/fetch.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/cli/fetch.py) during standard fetch operations.
- All recovered content includes **provenance metadata** and disclosure banners to maintain transparency.

## Frequently Asked Questions

### What triggers open-access recovery in Hyperresearch?

Recovery triggers in two scenarios: when `rescue_full_text` detects a fetch exception (HTTP 403/401 or connection errors) in [`src/hyperresearch/cli/fetch.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/cli/fetch.py), or when `recover_full_text` identifies content as "thin" (shorter than expected full-text) after a successful fetch. Both paths query the three OA APIs to find legal alternatives.

### Which open-access databases does Hyperresearch search?

The system searches **Unpaywall** (general coverage, requires contact email), **Europe PMC** (biomedical focus, no key required), and **CORE** (largest aggregator, requires API key). The `iter_oa_candidates` function queries them sequentially in that order.

### How does Hyperresearch prevent security risks when fetching OA copies?

Before downloading any candidate, `check_oa_url` (lines 34‑50 in [`src/hyperresearch/core/oa.py`](https://github.com/jordan-gibbs/hyperresearch/blob/main/src/hyperresearch/core/oa.py)) parses the URL, enforces HTTP(S) protocols, and resolves the IP address to confirm it is public. This blocks private network ranges and prevents SSRF attacks against internal infrastructure.

### What is the difference between rescued and substituted notes?

**Rescued** notes occur when the original URL is completely inaccessible; the system reads nothing from the source and replaces it entirely with OA content. **Substituted** notes occur when the original provides a thin abstract; the system preserves the original URL as the source but appends a banner and metadata indicating the full text came from an open-access repository.