# How to Use the Wigolo Extract Tool for Structured Data Extraction (JSON Schema & Tables)

> Learn how to use the Wigolo extract tool to get structured data from HTML. Convert unstructured content into JSON Schema objects or markdown tables easily.

- Repository: [Towhid Khan/wigolo](https://github.com/KnockOutEZ/wigolo)
- Tags: how-to-guide
- Published: 2026-07-29

---

**TLDR:** The wigolo extract tool fetches web content, processes it through the Defuddle extractor, and converts unstructured HTML into validated JSON Schema objects or markdown tables via the pipeline defined in [`src/tools/extract.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/tools/extract.ts).

Wigolo is an open-source web scraping and data extraction framework. Its `extract` command provides a low-level interface for turning unstructured web pages into structured datasets using **JSON Schema** validation, making it ideal for building automated data pipelines that require tabular or schema-defined outputs.

## How the Extraction Pipeline Works

The extraction flow is orchestrated in [`src/tools/extract.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/tools/extract.ts) and follows a three-stage pipeline:

1. **Fetch Layer** – Retrieves raw HTML via HTTP GET requests handled by the internal fetch utility.
2. **Extraction Provider** – Delegates processing to [`src/providers/extract-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/extract-provider.ts), which invokes the configured extractor (Defuddle by default) and returns an `ExtractionResult` containing markdown, metadata, and optional schema data.
3. **Post-Processing** – Applies section isolation, character limits, and JSON Schema validation before formatting the final output.

### Fetching and Provider Layer

When you run `wigolo extract <url>`, the tool initializes the provider defined in [`src/providers/extract-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/extract-provider.ts). This abstraction allows swapping extractors via the `--extractor` flag or `WIGOLO_EXTRACTOR` environment variable. The default **Defuddle** extractor parses HTML into clean markdown while preserving semantic structure for table detection.

### Schema Validation and Output Formatting

If you supply a `--schema` path, the tool validates extracted data against the JSON Schema before returning it. Validation errors surface specific field mismatches. Supported output formats include:

- **json** – Validated JSON objects ready for API ingestion.
- **markdown** – Raw extracted text with preserved formatting.
- **csv** – Tabular exports (generated via piping to external tools like `jq`).

## Step-by-Step Usage Examples

### Extract Validated JSON Using a Schema

Create a schema file defining your expected structure:

```json
// report-schema.json
{
  "$schema": "http://json-schema.org/draft-07/schema#",
  "type": "object",
  "properties": {
    "title": { "type": "string" },
    "date": { "type": "string", "format": "date-time" },
    "summary": { "type": "string" }
  },
  "required": ["title", "date", "summary"]
}

```

Run the extraction:

```bash
wigolo extract https://news.example.com/latest \
  --schema ./report-schema.json \
  --format json

```

**Output:**

```json
{
  "title": "New AI Model Beats Humans",
  "date": "2024-06-15T08:00:00Z",
  "summary": "A breakthrough in language modeling..."
}

```

### Isolate Specific Sections

Use section extraction to target specific content blocks, implemented in [`src/extraction/markdown.js`](https://github.com/KnockOutEZ/wigolo/blob/main/src/extraction/markdown.js):

```bash
wigolo extract https://research.example.com/paper \
  --section "Key Findings" \
  --max-chars 1500 \
  --format markdown

```

This returns only the first 1500 characters of the "Key Findings" section.

### Convert to CSV Tables

Pipe JSON output through `jq` to generate CSV tables for data analysis:

```bash
wigolo extract https://data.example.com/report \
  --schema ./schemas/report-schema.json \
  --format json | jq -r '(.[0] | keys_unsorted) as $keys | $keys, map([.[ $keys[] ]])[] | @csv'

```

## Core Source Files

Understanding the implementation requires referencing these key files:

- **[[`src/tools/extract.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/tools/extract.ts)](https://github.com/KnockOutEZ/wigolo/blob/main/src/tools/extract.ts)** – CLI entry point that orchestrates the fetch → extract → output pipeline.
- **[[`src/providers/extract-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/extract-provider.ts)](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/extract-provider.ts)** – Abstraction layer for extractor services; handles Defuddle API calls and response normalization.
- **[[`src/extraction/markdown.js`](https://github.com/KnockOutEZ/wigolo/blob/main/src/extraction/markdown.js)](https://github.com/KnockOutEZ/wigolo/blob/main/src/extraction/markdown.js)** – Utility functions for section isolation, character limits, and markdown post-processing.
- **[[`docs/tools.md`](https://github.com/KnockOutEZ/wigolo/blob/main/docs/tools.md)](https://github.com/KnockOutEZ/wigolo/blob/main/docs/tools.md)** – Complete flag reference and usage examples for the extract tool.

## Summary

- The **wigolo extract tool** converts web pages into structured data via a pipeline in [`src/tools/extract.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/tools/extract.ts).
- **JSON Schema validation** ensures output conforms to expected structures, ideal for database ingestion.
- The **provider pattern** in [`src/providers/extract-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/extract-provider.ts) supports pluggable extractors (Defuddle default).
- **Section extraction** and **character limits** are handled in [`src/extraction/markdown.js`](https://github.com/KnockOutEZ/wigolo/blob/main/src/extraction/markdown.js) for precise data targeting.
- Results can be output as **JSON**, **markdown**, or piped into **CSV** for tabular analysis.

## Frequently Asked Questions

### What extractors does wigolo support?

By default, wigolo uses the **Defuddle** extractor. You can switch to alternative extractors like `readability` or `goose` using the `--extractor` flag or by setting the `WIGOLO_EXTRACTOR` environment variable.

### How do I validate extracted data against a custom JSON Schema?

Pass the schema file path using the `--schema` flag. The tool validates extracted data before output and reports specific field errors if validation fails, helping you refine either the schema or extraction parameters.

### Can I extract specific sections instead of the full page?

Yes. Use the `--section` flag to target specific content blocks, and `--max-chars` to limit output length. These filters are implemented in [`src/extraction/markdown.js`](https://github.com/KnockOutEZ/wigolo/blob/main/src/extraction/markdown.js) and apply after the initial HTML extraction.

### Where is the extraction logic implemented in the source code?

The main orchestration resides in [`src/tools/extract.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/tools/extract.ts), which coordinates fetching and formatting. The actual extraction algorithms and section parsing utilities are located in [`src/extraction/markdown.js`](https://github.com/KnockOutEZ/wigolo/blob/main/src/extraction/markdown.js) and the provider abstraction in [`src/providers/extract-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/providers/extract-provider.ts).