How to Analyze and Generate Structured Data (JSON-LD) Using the Geo-SEO Toolkit

The Geo-SEO toolkit provides a complete pipeline for extracting existing JSON-LD from web pages and generating new structured data snippets using predefined schema templates.

The zubair-trabzada/geo-seo-claude repository delivers open-source utilities to analyze and generate structured data (JSON-LD) for modern SEO workflows. At its core, the scripts/fetch_page.py module scrapes live pages to extract embedded schema markup, while the schema/ directory provides battle-tested templates for creating compliant JSON-LD entities.

Extracting JSON-LD from Live Web Pages

The toolkit's page analyzer retrieves and parses structured data through a robust extraction pipeline implemented in scripts/fetch_page.py.

HTTP Request and HTML Parsing

The scraper initiates an HTTP request using realistic browser headers defined in DEFAULT_HEADERS, then constructs a navigable DOM using BeautifulSoup with the lxml parser. This approach ensures compatibility with modern JavaScript-heavy pages while maintaining parsing speed.

JSON-LD Detection and Extraction

Before any destructive DOM cleanup occurs, the scraper searches for <script type="application/ld+json"> blocks. For each discovered block, the tool executes json.loads(script.string) to parse the JavaScript content into native Python dictionaries. Successfully parsed objects append to result["structured_data"], creating a list of structured data entities ready for downstream processing.

from scripts.fetch_page import fetch_page

url = "https://example.com"
analysis = fetch_page(url)

# Access the list of parsed JSON-LD objects

json_ld_blocks = analysis["structured_data"]
for block in json_ld_blocks:
    print(block)  # Python dict representation

Error Handling for Malformed Data

The extraction logic includes defensive programming against corrupted schema markup. When json.loads() encounters invalid JSON or TypeError exceptions arise from null values, the catcher records these failures in result["errors"] without crashing the entire analysis pipeline.

Command-line users can inspect extracted blocks immediately using jq:

python scripts/fetch_page.py https://example.com page | jq '.structured_data'

Generating Structured Data from Templates

Beyond extraction, the toolkit simplifies creation of valid JSON-LD through predefined schema templates stored in the schema/ directory.

Working with Schema Templates

The repository includes ready-made templates such as schema/organization.json and schema/product-ecommerce.json. These files provide compliant starting points for common entity types including Organizations, Products, and Software applications.

import json
from pathlib import Path

# Load a schema template

template_path = Path("schema/organization.json")
template = json.loads(template_path.read_text())

# Fill in dynamic values

template["name"] = "Acme Corp"
template["url"] = "https://acme.example"
template["sameAs"] = [
    "https://en.wikipedia.org/wiki/Acme_Corp",
    "https://www.linkedin.com/company/acme"
]

# Serialize for embedding in HTML

json_ld = json.dumps(template, indent=2)
print(json_ld)   # Paste into <script type="application/ld+json">...</script>

Validating JSON-LD Output

Before deploying generated markup, verify serializability to prevent runtime errors:

import json

for block in json_ld_blocks:
    try:
        json.dumps(block)   # raises TypeError if not serializable

    except (TypeError, ValueError) as e:
        print("Invalid block:", e)

Summary

  • Analyze and generate structured data (JSON-LD) using the unified scripts/fetch_page.py utility and schema templates in zubair-trabzada/geo-seo-claude.
  • The extraction pipeline captures <script type="application/ld+json"> blocks before DOM manipulation, storing results in analysis["structured_data"].
  • Error handling captures json.JSONDecodeError and TypeError exceptions, logging issues to analysis["errors"] without stopping execution.
  • Predefined templates in schema/organization.json and schema/product-ecommerce.json accelerate compliant markup generation.
  • Integration tests in tests/test_fetch_page_ssr.py verify extraction accuracy across different page architectures.

Frequently Asked Questions

How does the Geo-SEO toolkit handle pages with multiple JSON-LD blocks?

The scraper captures every <script type="application/ld+json"> element found in the raw HTML and appends each successfully parsed object to the structured_data list. This design supports pages containing multiple entities—such as an Organization schema combined with individual Product schemas—returning them as distinct dictionary elements within the results array.

What happens when the scraper encounters malformed JSON-LD?

Invalid schema markup triggers exception handling that catches both json.JSONDecodeError and TypeError exceptions. Rather than failing completely, the tool records the error details in result["errors"] and continues processing remaining blocks. This ensures partial data recovery from pages containing mixed valid and invalid structured data.

Can I customize the predefined schema templates for specific industries?

Yes. The JSON files in the schema/ directory—including organization.json and product-ecommerce.json—serve as mutable templates. Load them via json.loads(), modify specific fields like name, url, or custom properties, then serialize back to HTML-ready strings. The skills/geo-schema/SKILL.md documentation provides guidance on extending these templates for specialized verticals.

Where is the JSON-LD extraction logic tested?

Unit tests verifying JSON-LD detection and parsing reside in tests/test_fetch_page_ssr.py. These tests validate that the fetch_page() function correctly identifies script tags, handles encoding variations, and properly structures the output dictionary containing both structured_data and errors arrays.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →