How to Use the Wigolo Extract Tool for Structured Data Extraction (JSON Schema & Tables)
TLDR: The wigolo extract tool fetches web content, processes it through the Defuddle extractor, and converts unstructured HTML into validated JSON Schema objects or markdown tables via the pipeline defined in src/tools/extract.ts.
Wigolo is an open-source web scraping and data extraction framework. Its extract command provides a low-level interface for turning unstructured web pages into structured datasets using JSON Schema validation, making it ideal for building automated data pipelines that require tabular or schema-defined outputs.
How the Extraction Pipeline Works
The extraction flow is orchestrated in src/tools/extract.ts and follows a three-stage pipeline:
- Fetch Layer – Retrieves raw HTML via HTTP GET requests handled by the internal fetch utility.
- Extraction Provider – Delegates processing to
src/providers/extract-provider.ts, which invokes the configured extractor (Defuddle by default) and returns anExtractionResultcontaining markdown, metadata, and optional schema data. - Post-Processing – Applies section isolation, character limits, and JSON Schema validation before formatting the final output.
Fetching and Provider Layer
When you run wigolo extract <url>, the tool initializes the provider defined in src/providers/extract-provider.ts. This abstraction allows swapping extractors via the --extractor flag or WIGOLO_EXTRACTOR environment variable. The default Defuddle extractor parses HTML into clean markdown while preserving semantic structure for table detection.
Schema Validation and Output Formatting
If you supply a --schema path, the tool validates extracted data against the JSON Schema before returning it. Validation errors surface specific field mismatches. Supported output formats include:
- json – Validated JSON objects ready for API ingestion.
- markdown – Raw extracted text with preserved formatting.
- csv – Tabular exports (generated via piping to external tools like
jq).
Step-by-Step Usage Examples
Extract Validated JSON Using a Schema
Create a schema file defining your expected structure:
// report-schema.json
{
"$schema": "http://json-schema.org/draft-07/schema#",
"type": "object",
"properties": {
"title": { "type": "string" },
"date": { "type": "string", "format": "date-time" },
"summary": { "type": "string" }
},
"required": ["title", "date", "summary"]
}
Run the extraction:
wigolo extract https://news.example.com/latest \
--schema ./report-schema.json \
--format json
Output:
{
"title": "New AI Model Beats Humans",
"date": "2024-06-15T08:00:00Z",
"summary": "A breakthrough in language modeling..."
}
Isolate Specific Sections
Use section extraction to target specific content blocks, implemented in src/extraction/markdown.js:
wigolo extract https://research.example.com/paper \
--section "Key Findings" \
--max-chars 1500 \
--format markdown
This returns only the first 1500 characters of the "Key Findings" section.
Convert to CSV Tables
Pipe JSON output through jq to generate CSV tables for data analysis:
wigolo extract https://data.example.com/report \
--schema ./schemas/report-schema.json \
--format json | jq -r '(.[0] | keys_unsorted) as $keys | $keys, map([.[ $keys[] ]])[] | @csv'
Core Source Files
Understanding the implementation requires referencing these key files:
- [
src/tools/extract.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/tools/extract.ts) – CLI entry point that orchestrates the fetch → extract → output pipeline. - [
src/providers/extract-provider.ts](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/extract-provider.ts) – Abstraction layer for extractor services; handles Defuddle API calls and response normalization. - [
src/extraction/markdown.js](https://github.com/KnockOutEZ/wigolo/blob/main/src/extraction/markdown.js) – Utility functions for section isolation, character limits, and markdown post-processing. - [
docs/tools.md](https://github.com/KnockOutEZ/wigolo/blob/main/docs/tools.md) – Complete flag reference and usage examples for the extract tool.
Summary
- The wigolo extract tool converts web pages into structured data via a pipeline in
src/tools/extract.ts. - JSON Schema validation ensures output conforms to expected structures, ideal for database ingestion.
- The provider pattern in
src/providers/extract-provider.tssupports pluggable extractors (Defuddle default). - Section extraction and character limits are handled in
src/extraction/markdown.jsfor precise data targeting. - Results can be output as JSON, markdown, or piped into CSV for tabular analysis.
Frequently Asked Questions
What extractors does wigolo support?
By default, wigolo uses the Defuddle extractor. You can switch to alternative extractors like readability or goose using the --extractor flag or by setting the WIGOLO_EXTRACTOR environment variable.
How do I validate extracted data against a custom JSON Schema?
Pass the schema file path using the --schema flag. The tool validates extracted data before output and reports specific field errors if validation fails, helping you refine either the schema or extraction parameters.
Can I extract specific sections instead of the full page?
Yes. Use the --section flag to target specific content blocks, and --max-chars to limit output length. These filters are implemented in src/extraction/markdown.js and apply after the initial HTML extraction.
Where is the extraction logic implemented in the source code?
The main orchestration resides in src/tools/extract.ts, which coordinates fetching and formatting. The actual extraction algorithms and section parsing utilities are located in src/extraction/markdown.js and the provider abstraction in src/providers/extract-provider.ts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →