How Catalogue Ontologies Are Compiled in Microsoft Ontology-Playground

Microsoft Ontology-Playground compiles catalogue ontologies at build time using the scripts/compile-catalogue.ts TypeScript compiler, which validates metadata, parses RDF/OWL files, performs round-trip verification, and outputs a single public/catalogue.json artifact consumed by the frontend.

The microsoft/Ontology-Playground repository manages ontology definitions across three tiers—official, community, and external—through a rigorous build pipeline that ensures data integrity and metadata consistency. Understanding how catalogue ontologies are compiled during the build process is essential for contributors maintaining ontology definitions and developers extending the platform's validation logic.

The 8-Stage Build Pipeline

The compilation script orchestrates a sequential pipeline that transforms raw ontology directories into a validated JSON catalogue. Each stage enforces specific constraints before admitting an entry to the final output.

Stage 1: Directory Discovery and Security Validation

The discoverOntologyDirs() function (lines 70-88 in scripts/compile-catalogue.ts) walks the catalogue/ folder to locate candidate ontologies. The script recognizes three distinct directory layouts:

  • Official tier: Direct subdirectories of catalogue/official/
  • Community tier: Nested paths following catalogue/community/<user>/<slug>/
  • External tier: Nested paths following catalogue/external/<source>/<slug>/

Security checks reject directory names containing non-slug characters or symbolic links, preventing path traversal attacks and ambiguous identifier generation.

Stage 2: Metadata Schema Validation

Each ontology folder must contain a metadata.json file processed by validateMetadata() (lines 45-60). The validator enforces required fields—name, description, and category—and verifies that the category value belongs to a predefined whitelist. Missing required fields or invalid categories halt the entry's inclusion immediately.

Stage 3: RDF/OWL Parsing and Model Extraction

The compiler identifies the first .rdf or .owl file in the directory and invokes parseRDF() from src/lib/rdf/parser.ts (lines 145-267). This function parses the XML into an internal Ontology object and extracts any DataBinding elements defined within the source document, preparing the structured data for downstream validation.

Stage 4: Round-Trip Serialization Verification

To guarantee data integrity, the pipeline serializes the parsed Ontology object back to RDF format using serializeToRDF() from src/lib/rdf/serializer.ts, then re-parses the output (lines 71-79 in compile-catalogue.ts). This round-trip check verifies that the parser/serializer pair operates without data loss, ensuring that runtime consumers receive semantically identical content.

Stage 5: Style and Naming Convention Checks

The parsed ontology passes through validateOntologyStyle() (lines 81-95), implemented in scripts/style-validator.ts. This stage enforces naming conventions, spelling rules, and style guidelines specific to the Ontology-Playground ecosystem. Any error-level style issue prevents the entry from being included in the final catalogue.

Stage 6: Entry ID Generation and Aggregation

For validated entries, the compiler generates a stable identifier derived from the tier and relative path—formatted as <tier>/<slug> or <tier>/<user>/<slug> (lines 97-110). The resulting catalogue entry aggregates this ID, the validated metadata, the parsed Ontology object, and extracted DataBinding elements into a cohesive structure.

Stage 7: JSON Output Generation

After processing all discovered directories, the script writes a single JSON document to public/catalogue.json via writeFileSync() (lines 30-33). The output includes a Unix timestamp, total entry count, and the complete array of compiled entries, optimized for direct consumption by the frontend application.

Stage 8: Build Integration and Bundling

The catalogue:build npm script, defined in package.json (lines 7-11), executes the compiler before the main application build. The generated public/catalogue.json is bundled by Vite and served as a static asset, ensuring zero runtime overhead for ontology parsing.

Running the Catalogue Compiler

Execute the compilation pipeline manually using npm or invoke the TypeScript script directly:


# From the repository root

npm run catalogue:build

# Equivalent direct invocation

npx tsx scripts/compile-catalogue.ts

The script outputs a checkmark (✔) for each successfully processed entry and terminates by writing public/catalogue.json.

Programmatic Validation

Reuse the build pipeline's validation logic in custom scripts to verify ontologies before submission:

import { parseRDF } from '../src/lib/rdf/parser';
import { serializeToRDF } from '../src/lib/rdf/serializer';
import { validateOntologyStyle } from '../scripts/style-validator';

// Assume `rdfString` contains the ontology XML
const { ontology, bindings } = parseRDF(rdfString);

// Check style conformance
const styleIssues = validateOntologyStyle(ontology);
if (styleIssues.some(i => i.severity === 'error')) {
  throw new Error('Style validation failed');
}

// Verify round-trip integrity
const roundTrip = serializeToRDF(ontology, bindings);
parseRDF(roundTrip); // Throws if serialization introduced data loss

Loading Compiled Ontologies in the UI

The frontend retrieves the pre-compiled catalogue via HTTP request to minimize client-side processing:

import { useEffect, useState } from 'react';

interface CatalogueEntry {
  id: string;
  name: string;
  description: string;
  category: string;
}

export function useCatalogue() {
  const [entries, setEntries] = useState<CatalogueEntry[]>([]);
  
  useEffect(() => {
    fetch('/catalogue.json')
      .then(r => r.json())
      .then((data: { entries: CatalogueEntry[]; timestamp: number }) => 
        setEntries(data.entries)
      );
  }, []);
  
  return entries;
}

Key Files and Implementation Details

The compilation process coordinates these critical modules:

Summary

  • Build-time compilation ensures catalogue ontologies are validated, parsed, and optimized before deployment.
  • Security hardening includes slug-only directory names and symlink rejection to prevent traversal attacks.
  • Data integrity is guaranteed through mandatory round-trip RDF serialization and re-parsing.
  • Style enforcement via scripts/style-validator.ts maintains consistent naming across all three tiers.
  • Runtime performance benefits from serving a single pre-compiled JSON file rather than parsing XML client-side.

Frequently Asked Questions

What triggers the catalogue compilation in Ontology-Playground?

The compilation runs automatically when executing npm run build because the catalogue:build script is configured as a prerequisite in package.json (lines 7-11). Contributors can also trigger it independently using npm run catalogue:build to generate public/catalogue.json without performing a full application build.

How does the build process validate ontology metadata?

The validateMetadata() function (lines 45-60 in scripts/compile-catalogue.ts) requires every ontology folder to contain a metadata.json file with name, description, and category fields. The category value is checked against a predefined whitelist, and missing fields cause immediate rejection of the entry.

Why is round-trip RDF verification performed during the build?

The pipeline serializes the parsed Ontology object back to XML and re-parses it (lines 71-79) to verify that src/lib/rdf/parser.ts and src/lib/rdf/serializer.ts form a lossless pair. This ensures that ontologies distributed to users match the semantic content of the source files exactly.

Where is the compiled catalogue stored and how is it accessed?

The compiler writes output to public/catalogue.json (lines 30-33), which Vite bundles into the application distribution. At runtime, the frontend fetches this file via standard HTTP requests to /catalogue.json, receiving a JSON object containing a timestamp, entry count, and the complete array of ontology definitions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →