# How to Provide Your Own Literature Corpus to the ARS Pipeline via the Material Passport

> Learn how to provide your own literature corpus to the ARS pipeline using the Material Passport YAML. Inject custom data for enhanced research skills.

- Repository: [Edward Cheng-I Wu/academic-research-skills](https://github.com/Imbad0202/academic-research-skills)
- Tags: how-to-guide
- Published: 2026-05-13

---

**Yes, you can inject a custom literature corpus into the Academic Research Skills (ARS) pipeline by populating the optional `literature_corpus[]` array in the Material Passport YAML file, which Phase 1 agents consume before falling back to external database searches.**

The Academic Research Skills (ARS) suite from Imbad0202/academic-research-skills uses a cross-stage YAML ledger called the **Material Passport** to persist metadata between Claude Code sessions. Since version v3.6.4, this passport supports an optional top-level field that allows you to provide your own literature corpus to the ARS pipeline via the Material Passport, enabling a corpus-first workflow where your pre-curated papers take precedence over external database queries.

## Understanding the Material Passport Schema

The Material Passport (Schema 9) functions as the single input port for user-provided bibliographic data. According to the schema definition in [`shared/contracts/passport/literature_corpus_entry.schema.json`](https://github.com/Imbad0202/academic-research-skills/blob/main/shared/contracts/passport/literature_corpus_entry.schema.json), each entry in the `literature_corpus[]` array must follow a CSL-JSON-compatible structure.

Required fields include:

- `citation_key`: A unique identifier for the entry
- `authors`: An array of author objects with `family` and `given` name fields
- `year`: The publication year as an integer
- `title`: The work's title
- `source_pointer`: A URI pointing to the actual file (e.g., `file:///home/user/papers/example.pdf`)

Optional private fields include `abstract` and `user_notes`, which the **Phase 1 literature consumers**—specifically the `bibliography_agent` in *deep-research* and the `literature_strategist_agent` in *academic-paper*—use for pre-screening without mutating the original corpus.

## How the Pipeline Consumes Your Custom Corpus

When you provide your own literature corpus to the ARS pipeline via the Material Passport, the orchestrator detects the non-empty `literature_corpus[]` field at runtime and switches to a **corpus-first → search-fills-gap** workflow.

This consumption protocol follows four architectural constraints documented in [`academic-pipeline/references/literature_corpus_consumers.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/academic-pipeline/references/literature_corpus_consumers.md):

1. **No silent skips**: Every entry must be acknowledged or explicitly rejected with cause
2. **No mutation**: Agents cannot modify the corpus array in-place
3. **Graceful fallback**: Parse failures trigger warnings, not pipeline termination
4. **Strict pre-screening**: Agents must validate entries against the schema before use

If the field is absent or empty, the agents automatically fall back to the default external-database-only search mode.

## Methods to Populate the literature_corpus Field

### Manual YAML Entry

You can manually construct a [`passport.yaml`](https://github.com/Imbad0202/academic-research-skills/blob/main/passport.yaml) file that includes the literature corpus array. This approach works well for small, curated collections where you need precise control over metadata.

```yaml

# passport.yaml

literature_corpus:
  - citation_key: "smith2024deep"
    authors:
      - family: "Smith"
        given: "Jane"
    year: 2024
    title: "Deep Learning for Systematic Review"
    source_pointer: "file:///home/user/papers/smith2024.pdf"
    abstract: |
      This paper proposes a meta-learning approach to automate systematic reviews...
    user_notes: "Important for methods section."

```

### Using the Folder Scan Adapter

For bulk ingestion of PDF directories, the repository provides [`scripts/adapters/folder_scan.py`](https://github.com/Imbad0202/academic-research-skills/blob/main/scripts/adapters/folder_scan.py), a reference implementation that extracts metadata and populates the passport automatically.

Execute the adapter from the repository root:

```bash
python -m scripts.adapters.folder_scan \
  --root /path/to/pdf/folder \
  --output passport.yaml

```

This adapter walks the specified directory, extracts PDF metadata, builds CSL-JSON style entries, and appends them to the `literature_corpus[]` array. Additional shipped adapters include [`zotero.py`](https://github.com/Imbad0202/academic-research-skills/blob/main/zotero.py) and [`obsidian.py`](https://github.com/Imbad0202/academic-research-skills/blob/main/obsidian.py), detailed in [`scripts/adapters/README.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/scripts/adapters/README.md).

### Programmatic Generation with Python

For dynamic workflows, you can append CSL-JSON records to an existing passport using Python:

```python
import json
import yaml
from pathlib import Path

def add_to_passport(passport_path: Path, csl_record: dict) -> None:
    """Append a CSL-JSON record to the `literature_corpus` field of a Material Passport."""
    passport = yaml.safe_load(passport_path.read_text()) or {}
    entry = {
        "citation_key": csl_record["id"],
        "authors": csl_record["author"],
        "year": int(csl_record["issued"]["date-parts"][0][0]),
        "title": csl_record["title"],
        "source_pointer": csl_record.get("URL", "file:///unknown"),
        "abstract": csl_record.get("abstract", ""),
        "user_notes": "",
    }
    passport.setdefault("literature_corpus", []).append(entry)
    passport_path.write_text(yaml.safe_dump(passport, sort_keys=False))

# Example usage

csl = json.loads(Path("smith2024.csl.json").read_text())
add_to_passport(Path("passport.yaml"), csl)

```

## Validating Your Corpus Before Execution

The repository includes [`scripts/check_literature_corpus_schema.py`](https://github.com/Imbad0202/academic-research-skills/blob/main/scripts/check_literature_corpus_schema.py), a CI-style validator that ensures schema compliance before the pipeline runs. This validator checks that all entries in `literature_corpus[]` conform to the JSON Schema defined in [`shared/contracts/passport/literature_corpus_entry.schema.json`](https://github.com/Imbad0202/academic-research-skills/blob/main/shared/contracts/passport/literature_corpus_entry.schema.json).

Run the validator manually:

```bash
python -m scripts.check_literature_corpus_schema \
  --passport passport.yaml

```

The command exits with status 0 if the passport is valid; any schema violations are printed to STDERR. This script is also invoked automatically by the GitHub Actions workflow [`pytest.yml`](https://github.com/Imbad0202/academic-research-skills/blob/main/pytest.yml) to prevent malformed corpuses from entering the pipeline.

## Summary

- **The Material Passport** ([`passport.yaml`](https://github.com/Imbad0202/academic-research-skills/blob/main/passport.yaml)) serves as the sole input port for custom literature corpuses in the ARS pipeline
- **The optional `literature_corpus[]` field** triggers a corpus-first workflow when present, causing Phase 1 agents to prioritize your entries over external searches
- **Schema compliance** is enforced by [`shared/contracts/passport/literature_corpus_entry.schema.json`](https://github.com/Imbad0202/academic-research-skills/blob/main/shared/contracts/passport/literature_corpus_entry.schema.json) and validated via [`scripts/check_literature_corpus_schema.py`](https://github.com/Imbad0202/academic-research-skills/blob/main/scripts/check_literature_corpus_schema.py)
- **Three ingestion methods** exist: manual YAML editing, reference adapters like [`scripts/adapters/folder_scan.py`](https://github.com/Imbad0202/academic-research-skills/blob/main/scripts/adapters/folder_scan.py), and custom Python scripts
- **No runtime flags** are required; the pipeline automatically detects and consumes the field when populated

## Frequently Asked Questions

### What file format does the literature corpus use?

The `literature_corpus[]` field uses a CSL-JSON-compatible structure defined in [`shared/contracts/passport/literature_corpus_entry.schema.json`](https://github.com/Imbad0202/academic-research-skills/blob/main/shared/contracts/passport/literature_corpus_entry.schema.json). Each entry requires `citation_key`, `authors`, `year`, `title`, and `source_pointer` fields, with optional `abstract` and `user_notes` fields for private metadata.

### Can I mix my own corpus with external database searches?

Yes. When you provide your own literature corpus to the ARS pipeline via the Material Passport, the Phase 1 agents operate in **corpus-first → search-fills-gap** mode. They consume your pre-curated entries first, then use external database searches only to fill gaps identified in the literature review.

### How do I validate my Material Passport before running the pipeline?

Use the CI-style validator: `python -m scripts.check_literature_corpus_schema --passport passport.yaml`. This script enforces strict schema compliance and is automatically run by the [`pytest.yml`](https://github.com/Imbad0202/academic-research-skills/blob/main/pytest.yml) GitHub Actions workflow, ensuring that invalid entries are caught before the pipeline executes.

### Which agents actually read the literature_corpus field?

The **Phase 1 literature consumers** read this field: specifically, the `bibliography_agent` in the *deep-research* module and the `literature_strategist_agent` in the *academic-paper* module. These agents follow the four Iron Rules documented in [`academic-pipeline/references/literature_corpus_consumers.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/academic-pipeline/references/literature_corpus_consumers.md) to ensure safe, immutable consumption of your corpus.