# How the cards.json Data File Is Managed and Updated in the Pokémon TCG Pocket Collection Tracker

> Discover how the cards.json data file is automatically generated and updated. Learn about the Node.js scraper that extracts Pokémon TCG metadata from external sources.

- Repository: [Marcel Panse/tcg-pocket-collection-tracker](https://github.com/marcelpanse/tcg-pocket-collection-tracker)
- Tags: internals
- Published: 2026-03-06

---

**The [`cards.json`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/cards.json) file is automatically generated by a Node.js scraper ([`scripts/scraper.ts`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/scripts/scraper.ts)) that crawls pocket.limitlesstcg.com, extracts card metadata, downloads images, and writes normalized JSON to [`frontend/assets/cards.json`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/frontend/assets/cards.json) for runtime consumption by the React frontend.**

The [`cards.json`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/cards.json) file serves as the master data catalogue for the **Pokémon TCG Pocket Collection Tracker** open-source project. Located in the `marcelpanse/tcg-pocket-collection-tracker` repository, this file powers the entire card collection interface. It is never edited manually; instead, an automated pipeline scrapes external Pokémon TCG sources, normalizes the data, and hydrates the frontend with the latest card information.

## The Scraping Pipeline Architecture

The scraper logic resides in [`scripts/scraper.ts`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/scripts/scraper.ts) and orchestrates an eight-step transformation from raw HTML to structured JSON.

### Loading Expansion Metadata

The script begins by importing the canonical expansion list from [`frontend/src/lib/CardsDB.ts`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/frontend/src/lib/CardsDB.ts) to determine which sets to scrape. This ensures the scraper stays synchronized with the frontend's supported expansions.

### Discovering Card URLs

For each expansion, the script constructs a target URL using the pattern `https://pocket.limitlesstcg.com/cards/${expansion.id}`. Using **cheerio**, it parses the card grid HTML and extracts individual card page URLs from anchor tags within the `.card-search-grid` CSS selector, as implemented at lines 19-23 of [`scraper.ts`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/scraper.ts).

### Parallel Fetching and Parsing

The `processBatch()` function processes up to 10 card URLs concurrently (lines 44-56). For each page, the `extractCardInfo` function parses the HTML to extract:

- **Visual assets** (downloaded once to `frontend/public/images/en-US/`)
- **Gameplay statistics** (HP, energy type, attacks, abilities, weakness, retreat cost)
- **Rarity classifications** (including expansion-specific manual overrides)
- **Alternate versions** and deterministic internal IDs calculated via [`scripts/encoder.ts`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/scripts/encoder.ts)

### Applying Data Overrides

A hard-coded `deckbuildingOverrides` array (lines 5-6) reorders specific cards' `alternate_versions` arrays to ensure preferred variants appear first in the dataset, optimizing the deck-building experience.

### Sorting and Output Generation

After processing all expansions, the temporary card array is sorted by expansion ID and numeric card number. The script deduplicates alternate IDs and writes the final, pretty-printed JSON to [`frontend/assets/cards.json`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/frontend/assets/cards.json) (lines 74-101).

## Running the Scraper and Updating Data

To regenerate the card database, execute the scraper command from the repository root:

```bash
pnpm install
pnpm scraper  # Executes scripts/scraper.ts via Node

```

This command re-downloads card pages and images, regenerates [`frontend/assets/cards.json`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/frontend/assets/cards.json), and completely overwrites the previous dataset with the latest upstream data.

### Synchronizing Translations

After updating [`cards.json`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/cards.json), maintain translation parity by running:

```bash
node scripts/sync-card-translations.js

```

This utility (lines 42-109) scans the generated card list and inserts placeholder entries into [`pokemon_translations.json`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/pokemon_translations.json), [`tools_translations.json`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/tools_translations.json), and [`trainers_translations.json`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/trainers_translations.json) for supported locales (`en-US`, `es-ES`, `pt-BR`, `fr-FR`, `it-IT`).

## Frontend Integration via Dynamic Import

The React frontend does not bundle the JSON at build time. Instead, [`frontend/src/lib/CardsDB.ts`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/frontend/src/lib/CardsDB.ts) loads the catalogue dynamically:

```typescript
const AllCardsJson = await import('../../assets/cards.json')
export const AllCards = AllCardsJson.default

```

The `AllCards` constant becomes the single source of truth for React Query hooks, filtering utilities, collection tracking, and deck-building features throughout the application.

## CI/CD Validation and Deployment

The repository includes [`.github/workflows/verify-hashes.yml`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/.github/workflows/verify-hashes.yml), which validates the cryptographic hash of [`cards.json`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/cards.json) during continuous integration. This ensures committed data files match the expected output of the scraping pipeline before the build pipeline bundles them into the deployed frontend.

## Summary

- The [`cards.json`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/cards.json) file is auto-generated by [`scripts/scraper.ts`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/scripts/scraper.ts), not maintained manually.
- Data is sourced from **pocket.limitlesstcg.com** using **cheerio** for HTML parsing.
- The pipeline fetches up to 10 cards concurrently, downloads images to `frontend/public/images/en-US/`, and calculates internal IDs via [`scripts/encoder.ts`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/scripts/encoder.ts).
- Manual `deckbuildingOverrides` adjust alternate version ordering for specific cards.
- The frontend consumes data via dynamic import in [`CardsDB.ts`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/CardsDB.ts).
- Translation files stay synchronized via [`scripts/sync-card-translations.js`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/scripts/sync-card-translations.js).
- GitHub Actions validates file integrity via [`verify-hashes.yml`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/verify-hashes.yml) before deployment.

## Frequently Asked Questions

### How often should I run the scraper?

Run `pnpm scraper` whenever new expansions drop on pocket.limitlesstcg.com or when card errata require updated metadata. The script is idempotent and regenerates the entire file from scratch, ensuring consistency with the upstream source.

### Can I manually edit cards.json?

Manual edits are strongly discouraged because the scraper overwrites this file completely during execution. Instead, modify the scraping logic in [`scripts/scraper.ts`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/scripts/scraper.ts) or apply overrides in the `deckbuildingOverrides` array to adjust specific card attributes.

### Why does the scraper download images separately instead of using external URLs?

The pipeline downloads card images to `frontend/public/images/en-US/` to ensure the frontend serves optimized local assets rather than hotlinking external URLs. This improves page load times, ensures availability if the upstream site changes, and allows for offline-first capabilities.

### What happens if the upstream site HTML structure changes?

If pocket.limitlesstcg.com modifies its DOM structure, the cheerio selectors in `extractCardInfo` (lines 30-104 of [`scripts/scraper.ts`](https://github.com/marcelpanse/tcg-pocket-collection-tracker/blob/main/scripts/scraper.ts)) will require updating to match the new CSS selectors. The scraper would fail to extract data or throw parsing errors until the selectors are adjusted to reflect the new HTML layout.