Understanding the Site Snapshots System for Published Content in Instatic

In Instatic, the site_snapshots system creates an immutable JSON representation of the entire site (SiteDocument) for each publish, linking every published content row to this snapshot and serving pages through a three-layer rendering pipeline that guarantees atomic publishes and fast, reliable reads.

The site_snapshots mechanism is the core of the publishing pipeline in the CoreBunch/Instatic repository. It provides a single immutable representation of the whole site that is created once per publish and then used by every layer that serves public pages. This design ensures strong isolation between draft edits and what visitors see, while enabling fast, cacheable rendering.

How the Site Snapshots System Stores Published Content

The storage layer consists of three interconnected components that together guarantee auditability and prevent partial publish states.

SQL Storage Layer

The site_snapshots table holds the JSON-encoded SiteDocument containing the full page tree, layout CSS, import-map, and all site metadata. The schema is defined in server/db/migrations-pg.ts (lines 587-610) for PostgreSQL and server/db/migrations-sqlite.ts (lines 528-551) for SQLite.

Each row represents a complete, frozen version of the site at a specific point in time. The content_hash column stored on the row enables cheap change-detection between publishes.

Content Row References

Every published data-row version points to the snapshot that was current when it was published via the data_row_versions.site_snapshot_id foreign key. This linkage is managed in server/repositories/publish.ts (lines 5-12), ensuring that historical content versions always reference the exact site structure that existed at their time of publication.

Atomic Insertion Strategy

During a publish transaction, the system inserts a new row into site_snapshots and then updates all data_row_versions to reference the new snapshot ID. This occurs in server/repositories/publish.ts approximately lines 236-258. Because this happens within a single transaction, readers never see a partially-published state—either the entire publish succeeds or the database remains unchanged.

Three-Layer Rendering Pipeline

The site_snapshots system powers a three-layer rendering pipeline that optimizes for different types of requests:

Layer A: Static Artefacts

When a publish finishes, server/publish/staticArtefact.ts writes fully baked HTML for each route to disk at uploads/published/current/<route>.html. This provides the fastest possible response for pages that have no request-dependent fragments, serving pure static files without touching the database or cache layers.

Layer B: In-Memory LRU Cache

For pages that need the snapshot, the public router first checks the Layer B cache implemented in server/publish/publishedSnapshotCache.ts. The cache key is (urlPath, queryString, publishVersion).

On a cache miss, the system loads the SiteDocument from site_snapshots, re-assembles a PublishedPageSnapshot, renders the HTML, and stores the result in the LRU cache. When a new publish occurs, bumpPublishVersion() clears the entire cache, ensuring the next request loads the fresh snapshot.

Layer C: Dynamic Hole Runtime

During the render walk, the publisher detects nodes that depend on the request (such as loop bindings or route.query values). These nodes are replaced by an <instatic-hole> placeholder. The client-side runtime in server/publish/holeRuntime.ts serves fragments on demand via the /_instatic/hole/:nodeId endpoint, allowing dynamic content without re-walking the entire tree.

Request Resolution Flow

When a visitor requests a public page, the resolution follows this path:

  1. Router entry — server/router.ts forwards public URLs to server/publish/publicRouter.ts.
  2. Row lookup — publicRouter.ts joins data_row_versions to site_snapshots to locate the correct SiteDocument.
  3. Snapshot-aware rendering — server/publish/publicRenderer.ts calls renderPublishedSnapshot(snapshot, …) which walks the immutable tree (no live mutations occur during rendering).
  4. Layer selection — If the rendered page has no dynamic holes, the result comes from the static artefact (Layer A). Otherwise, the HTML is cached in Layer B and any holes are served by Layer C.

Benefits of the Site Snapshots Architecture

  • Atomic publishes — The whole site freezes into a single JSON blob; readers never see partially-published states.
  • Fast reads — Layer B cache eliminates database work after the first request for a given URL/version combination.
  • Audit trail — Each publish creates a new site_snapshots row, allowing reconstruction of any historic public version.
  • Strong isolation — Draft edits remain in the live store and never affect the snapshot until a new publish bumps publishVersion.
  • Cache-friendly — The snapshot hash enables efficient change detection and cache invalidation.
  • Dynamic fragment support — Request-dependent nodes are lazily hydrated via Layer C without re-rendering the entire page.

Code Examples

Publishing a Site

import { publishSite } from '@core/publish';

// Inside the publish command
await publishSite({
  db,
  registry,
  // options like commitMessage, preview, etc.
});

Underlying steps in server/publish/publishSite.ts:

  1. Build the SiteDocument via src/core/persistence/serializeSite.ts.
  2. Insert a new row into site_snapshots.
  3. Update all data_row_versions to reference the new snapshot ID.
  4. Write static HTML artefacts via staticArtefact.ts.

Rendering a Public Page


# Direct HTTP request (no auth) – layer B cache will be used

curl https://example.com/blog/my-post

Server flow:

  1. publicRouter.ts → getPublishedPageBySlug() joins to site_snapshots.
  2. renderPublishedSnapshot(snapshot, { url, publishVersion }) returns HTML.
  3. If the page contains holes, the initial HTML contains <instatic-hole data-node-id="xyz">…</instatic-hole>.
  4. Browser fetches /_instatic/hole/xyz?publishVersion=123 → holeRuntime.ts renders the fragment on-demand.

Accessing Snapshots Programmatically

import { SiteAgentSnapshotSchema } from '@core/ai/tools/site/snapshot';
import { safeParseValue } from '@core/utils/safeParse';

// Inside an AI tool handler
const snapshot = await loadCurrentSiteSnapshot(db);
const parsed = safeParseValue(SiteAgentSnapshotSchema, snapshot);
if (!parsed.success) { 
  /* handle malformed snapshot */ 
}

The loadCurrentSiteSnapshot function resides in server/publish/publishedSnapshotCache.ts, using the same cache layer as the public router.

Key Source Files

File Purpose
server/repositories/publish.ts Inserts new site_snapshots rows and wires data_row_versions references.
server/publish/publishedSnapshotCache.ts LRU cache implementation for rendered HTML per publish version.
server/publish/publicRouter.ts Joins data rows to snapshots and selects entry templates.
server/publish/publicRenderer.ts Contains renderPublishedSnapshot and renderPublishedDataRowTemplate functions.
server/publish/staticArtefact.ts Generates static HTML files for Layer A.
server/publish/holeRuntime.ts Runtime handling of dynamic hole fragments (Layer C).
server/db/migrations-pg.ts & server/db/migrations-sqlite.ts Database schema for site_snapshots and foreign key constraints.
docs/features/publisher.md Architectural description of the three-layer pipeline.
docs/server.md Overview of public router delegation to the snapshot system.

Summary

  • The site_snapshots system stores an immutable JSON SiteDocument for each publish in the CoreBunch/Instatic repository.
  • Every published content row references its snapshot via data_row_versions.site_snapshot_id, ensuring historical accuracy.
  • Pages are served through a three-layer pipeline: static artefacts (Layer A), in-memory LRU cache (Layer B), and dynamic hole runtime (Layer C).
  • The architecture guarantees atomic publishes, fast cacheable reads, and complete auditability through immutable snapshots.

Frequently Asked Questions

What is stored in the site_snapshots row?

Each row contains a JSON-encoded SiteDocument that includes the full page tree, layout CSS, import-map, and all site metadata. The row also stores a content_hash for change detection and versioning. This immutable blob represents the entire site state at the moment of publish.

How does cache invalidation work when publishing new content?

When a new publish completes, the bumpPublishVersion() function clears the entire Layer B in-memory LRU cache. The next request for any URL triggers a fresh load from the site_snapshots table, ensuring visitors never see stale content. Static artefacts (Layer A) are overwritten during the publish process.

What is the difference between static artefacts and dynamic holes?

Static artefacts are fully baked HTML files written to disk during publish (Layer A) for pages with no request-dependent content. Dynamic holes (Layer C) are placeholder elements (<instatic-hole>) inserted during rendering for nodes that depend on request parameters like query strings or user-specific data. These holes are hydrated on-demand via the holeRuntime.ts endpoint.

How does the site_snapshots system ensure atomic publishes?

The system uses a database transaction that inserts the new site_snapshots row and updates all data_row_versions to reference it within a single atomic operation. This occurs in server/repositories/publish.ts (lines 236-258). If any step fails, the transaction rolls back, preventing partial publishes where some pages reference the new snapshot while others reference the old.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →