How to Use the Wigolo Extract Tool for Structured Data Extraction (JSON Schema & Tables)

TLDR: The wigolo extract tool fetches web content, processes it through the Defuddle extractor, and converts unstructured HTML into validated JSON Schema objects or markdown tables via the pipeline defined in src/tools/extract.ts.

Wigolo is an open-source web scraping and data extraction framework. Its extract command provides a low-level interface for turning unstructured web pages into structured datasets using JSON Schema validation, making it ideal for building automated data pipelines that require tabular or schema-defined outputs.

How the Extraction Pipeline Works

The extraction flow is orchestrated in src/tools/extract.ts and follows a three-stage pipeline:

  1. Fetch Layer – Retrieves raw HTML via HTTP GET requests handled by the internal fetch utility.
  2. Extraction Provider – Delegates processing to src/providers/extract-provider.ts, which invokes the configured extractor (Defuddle by default) and returns an ExtractionResult containing markdown, metadata, and optional schema data.
  3. Post-Processing – Applies section isolation, character limits, and JSON Schema validation before formatting the final output.

Fetching and Provider Layer

When you run wigolo extract <url>, the tool initializes the provider defined in src/providers/extract-provider.ts. This abstraction allows swapping extractors via the --extractor flag or WIGOLO_EXTRACTOR environment variable. The default Defuddle extractor parses HTML into clean markdown while preserving semantic structure for table detection.

Schema Validation and Output Formatting

If you supply a --schema path, the tool validates extracted data against the JSON Schema before returning it. Validation errors surface specific field mismatches. Supported output formats include:

  • json – Validated JSON objects ready for API ingestion.
  • markdown – Raw extracted text with preserved formatting.
  • csv – Tabular exports (generated via piping to external tools like jq).

Step-by-Step Usage Examples

Extract Validated JSON Using a Schema

Create a schema file defining your expected structure:

// report-schema.json
{
  "$schema": "http://json-schema.org/draft-07/schema#",
  "type": "object",
  "properties": {
    "title": { "type": "string" },
    "date": { "type": "string", "format": "date-time" },
    "summary": { "type": "string" }
  },
  "required": ["title", "date", "summary"]
}

Run the extraction:

wigolo extract https://news.example.com/latest \
  --schema ./report-schema.json \
  --format json

Output:

{
  "title": "New AI Model Beats Humans",
  "date": "2024-06-15T08:00:00Z",
  "summary": "A breakthrough in language modeling..."
}

Isolate Specific Sections

Use section extraction to target specific content blocks, implemented in src/extraction/markdown.js:

wigolo extract https://research.example.com/paper \
  --section "Key Findings" \
  --max-chars 1500 \
  --format markdown

This returns only the first 1500 characters of the "Key Findings" section.

Convert to CSV Tables

Pipe JSON output through jq to generate CSV tables for data analysis:

wigolo extract https://data.example.com/report \
  --schema ./schemas/report-schema.json \
  --format json | jq -r '(.[0] | keys_unsorted) as $keys | $keys, map([.[ $keys[] ]])[] | @csv'

Core Source Files

Understanding the implementation requires referencing these key files:

Summary

  • The wigolo extract tool converts web pages into structured data via a pipeline in src/tools/extract.ts.
  • JSON Schema validation ensures output conforms to expected structures, ideal for database ingestion.
  • The provider pattern in src/providers/extract-provider.ts supports pluggable extractors (Defuddle default).
  • Section extraction and character limits are handled in src/extraction/markdown.js for precise data targeting.
  • Results can be output as JSON, markdown, or piped into CSV for tabular analysis.

Frequently Asked Questions

What extractors does wigolo support?

By default, wigolo uses the Defuddle extractor. You can switch to alternative extractors like readability or goose using the --extractor flag or by setting the WIGOLO_EXTRACTOR environment variable.

How do I validate extracted data against a custom JSON Schema?

Pass the schema file path using the --schema flag. The tool validates extracted data before output and reports specific field errors if validation fails, helping you refine either the schema or extraction parameters.

Can I extract specific sections instead of the full page?

Yes. Use the --section flag to target specific content blocks, and --max-chars to limit output length. These filters are implemented in src/extraction/markdown.js and apply after the initial HTML extraction.

Where is the extraction logic implemented in the source code?

The main orchestration resides in src/tools/extract.ts, which coordinates fetching and formatting. The actual extraction algorithms and section parsing utilities are located in src/extraction/markdown.js and the provider abstraction in src/providers/extract-provider.ts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →