How to Provide Your Own Literature Corpus to the ARS Pipeline via the Material Passport
Yes, you can inject a custom literature corpus into the Academic Research Skills (ARS) pipeline by populating the optional literature_corpus[] array in the Material Passport YAML file, which Phase 1 agents consume before falling back to external database searches.
The Academic Research Skills (ARS) suite from Imbad0202/academic-research-skills uses a cross-stage YAML ledger called the Material Passport to persist metadata between Claude Code sessions. Since version v3.6.4, this passport supports an optional top-level field that allows you to provide your own literature corpus to the ARS pipeline via the Material Passport, enabling a corpus-first workflow where your pre-curated papers take precedence over external database queries.
Understanding the Material Passport Schema
The Material Passport (Schema 9) functions as the single input port for user-provided bibliographic data. According to the schema definition in shared/contracts/passport/literature_corpus_entry.schema.json, each entry in the literature_corpus[] array must follow a CSL-JSON-compatible structure.
Required fields include:
citation_key: A unique identifier for the entryauthors: An array of author objects withfamilyandgivenname fieldsyear: The publication year as an integertitle: The work's titlesource_pointer: A URI pointing to the actual file (e.g.,file:///home/user/papers/example.pdf)
Optional private fields include abstract and user_notes, which the Phase 1 literature consumers—specifically the bibliography_agent in deep-research and the literature_strategist_agent in academic-paper—use for pre-screening without mutating the original corpus.
How the Pipeline Consumes Your Custom Corpus
When you provide your own literature corpus to the ARS pipeline via the Material Passport, the orchestrator detects the non-empty literature_corpus[] field at runtime and switches to a corpus-first → search-fills-gap workflow.
This consumption protocol follows four architectural constraints documented in academic-pipeline/references/literature_corpus_consumers.md:
- No silent skips: Every entry must be acknowledged or explicitly rejected with cause
- No mutation: Agents cannot modify the corpus array in-place
- Graceful fallback: Parse failures trigger warnings, not pipeline termination
- Strict pre-screening: Agents must validate entries against the schema before use
If the field is absent or empty, the agents automatically fall back to the default external-database-only search mode.
Methods to Populate the literature_corpus Field
Manual YAML Entry
You can manually construct a passport.yaml file that includes the literature corpus array. This approach works well for small, curated collections where you need precise control over metadata.
# passport.yaml
literature_corpus:
- citation_key: "smith2024deep"
authors:
- family: "Smith"
given: "Jane"
year: 2024
title: "Deep Learning for Systematic Review"
source_pointer: "file:///home/user/papers/smith2024.pdf"
abstract: |
This paper proposes a meta-learning approach to automate systematic reviews...
user_notes: "Important for methods section."
Using the Folder Scan Adapter
For bulk ingestion of PDF directories, the repository provides scripts/adapters/folder_scan.py, a reference implementation that extracts metadata and populates the passport automatically.
Execute the adapter from the repository root:
python -m scripts.adapters.folder_scan \
--root /path/to/pdf/folder \
--output passport.yaml
This adapter walks the specified directory, extracts PDF metadata, builds CSL-JSON style entries, and appends them to the literature_corpus[] array. Additional shipped adapters include zotero.py and obsidian.py, detailed in scripts/adapters/README.md.
Programmatic Generation with Python
For dynamic workflows, you can append CSL-JSON records to an existing passport using Python:
import json
import yaml
from pathlib import Path
def add_to_passport(passport_path: Path, csl_record: dict) -> None:
"""Append a CSL-JSON record to the `literature_corpus` field of a Material Passport."""
passport = yaml.safe_load(passport_path.read_text()) or {}
entry = {
"citation_key": csl_record["id"],
"authors": csl_record["author"],
"year": int(csl_record["issued"]["date-parts"][0][0]),
"title": csl_record["title"],
"source_pointer": csl_record.get("URL", "file:///unknown"),
"abstract": csl_record.get("abstract", ""),
"user_notes": "",
}
passport.setdefault("literature_corpus", []).append(entry)
passport_path.write_text(yaml.safe_dump(passport, sort_keys=False))
# Example usage
csl = json.loads(Path("smith2024.csl.json").read_text())
add_to_passport(Path("passport.yaml"), csl)
Validating Your Corpus Before Execution
The repository includes scripts/check_literature_corpus_schema.py, a CI-style validator that ensures schema compliance before the pipeline runs. This validator checks that all entries in literature_corpus[] conform to the JSON Schema defined in shared/contracts/passport/literature_corpus_entry.schema.json.
Run the validator manually:
python -m scripts.check_literature_corpus_schema \
--passport passport.yaml
The command exits with status 0 if the passport is valid; any schema violations are printed to STDERR. This script is also invoked automatically by the GitHub Actions workflow pytest.yml to prevent malformed corpuses from entering the pipeline.
Summary
- The Material Passport (
passport.yaml) serves as the sole input port for custom literature corpuses in the ARS pipeline - The optional
literature_corpus[]field triggers a corpus-first workflow when present, causing Phase 1 agents to prioritize your entries over external searches - Schema compliance is enforced by
shared/contracts/passport/literature_corpus_entry.schema.jsonand validated viascripts/check_literature_corpus_schema.py - Three ingestion methods exist: manual YAML editing, reference adapters like
scripts/adapters/folder_scan.py, and custom Python scripts - No runtime flags are required; the pipeline automatically detects and consumes the field when populated
Frequently Asked Questions
What file format does the literature corpus use?
The literature_corpus[] field uses a CSL-JSON-compatible structure defined in shared/contracts/passport/literature_corpus_entry.schema.json. Each entry requires citation_key, authors, year, title, and source_pointer fields, with optional abstract and user_notes fields for private metadata.
Can I mix my own corpus with external database searches?
Yes. When you provide your own literature corpus to the ARS pipeline via the Material Passport, the Phase 1 agents operate in corpus-first → search-fills-gap mode. They consume your pre-curated entries first, then use external database searches only to fill gaps identified in the literature review.
How do I validate my Material Passport before running the pipeline?
Use the CI-style validator: python -m scripts.check_literature_corpus_schema --passport passport.yaml. This script enforces strict schema compliance and is automatically run by the pytest.yml GitHub Actions workflow, ensuring that invalid entries are caught before the pipeline executes.
Which agents actually read the literature_corpus field?
The Phase 1 literature consumers read this field: specifically, the bibliography_agent in the deep-research module and the literature_strategist_agent in the academic-paper module. These agents follow the four Iron Rules documented in academic-pipeline/references/literature_corpus_consumers.md to ensure safe, immutable consumption of your corpus.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →