How to Extract Clean Markdown from Web Pages Using Defuddle

The defuddle skill is a lightweight wrapper around the Defuddle CLI that converts web pages into clean, token-efficient markdown by stripping navigation, ads, and HTML boilerplate.

The obsidian-skills repository by kepano provides agent-compatible tools for knowledge management workflows. When you need to extract clean markdown from web pages using Defuddle, the defuddle skill offers a streamlined interface that parses HTML and returns only the essential content structure. This approach significantly reduces token consumption for downstream LLM processing compared to raw web fetching, as documented in the repository's README.md and skills/defuddle/SKILL.md.

Understanding the Defuddle Skill Architecture

Each skill in the repository adheres to the Agent Skills specification and is auto-discovered by hosts like Claude Code, Codex CLI, or OpenCode through a SKILL.md file. The defuddle skill definition resides in skills/defuddle/SKILL.md, where it declares the skill name, description, and exact CLI commands for the host to execute. This declarative approach allows agents to invoke defuddle parse commands without hard-coded logic, simply by reading the skill definition file.

Installing the Defuddle CLI

Before invoking the skill, the Defuddle CLI must be present on the host machine. According to skills/defuddle/SKILL.md, install the tool globally via npm:

npm install -g defuddle

Once installed, the skill becomes immediately available because the host forwards user-provided URLs directly to this CLI. No additional configuration files are required within the agent environment.

Extracting Clean Markdown from Web Pages

The primary function converts noisy HTML into structured markdown. The agent resolves the skill and executes the command documented in the Usage section of skills/defuddle/SKILL.md.

Basic Markdown Extraction

To retrieve clean markdown from any public URL, use the --md flag:

defuddle parse https://example.com/article --md

This strips navigation bars, advertisements, and scripting elements, returning only the article content in markdown format.

Saving Output to Files

For persistent storage or subsequent processing by other Obsidian skills, redirect output to a file:

defuddle parse https://example.com/article --md -o article.md

Retrieving Specific Metadata

When you need only specific fields rather than full content, use the -p flag with property names:

defuddle parse https://example.com/article -p title

Available properties include title, description, and domain.

Output Format Options

The skill supports multiple output modes as documented in the Output formats table within skills/defuddle/SKILL.md:

  • --md – Returns clean markdown (default and most token-efficient for LLM consumption)
  • --json – Returns JSON containing both raw HTML and generated markdown
  • (no flag) – Returns raw, unprocessed HTML
  • -p <name> – Returns specific metadata values

Integration with Agent Workflows

Because the defuddle skill returns pure markdown, it integrates seamlessly with other skills in the obsidian-skills ecosystem. The output can be piped directly into obsidian-markdown or similar skills to create structured notes without intermediate parsing steps. This architecture reduces the number of tokens the LLM must process by eliminating HTML boilerplate before content reaches the reasoning engine.

Summary

  • The defuddle skill in skills/defuddle/SKILL.md wraps the Defuddle CLI for agent-compatible web parsing
  • Install the CLI globally using npm install -g defuddle
  • Use defuddle parse <url> --md to extract clean, token-efficient markdown
  • Choose alternative output formats with --json or metadata extraction with -p <property>
  • The skill integrates with other Obsidian tools by returning pure markdown suitable for direct note creation

Frequently Asked Questions

What is Defuddle and how does it differ from standard web scraping?

Defuddle is a CLI tool that extracts semantic content from HTML, specifically designed to remove navigation, ads, and decorative elements that standard scraping tools often retain. Unlike generic fetch utilities that return raw HTML, Defuddle outputs clean markdown that is immediately usable by LLMs without preprocessing, reducing token counts by 60-80% on typical web pages.

How do I install the defuddle skill in my agent environment?

The skill requires only the Defuddle CLI to be installed on the host machine via npm install -g defuddle. The skill definition in skills/defuddle/SKILL.md is automatically discovered by agents following the Agent Skills specification, requiring no additional configuration or registration steps.

Can I extract specific metadata instead of full markdown content?

Yes. Use the -p flag followed by the property name to retrieve specific fields. For example, defuddle parse <url> -p title returns only the page title. Supported properties include title, description, and domain as documented in the skill's output formats table in skills/defuddle/SKILL.md.

How does defuddle improve LLM token efficiency compared to raw HTML?

By parsing HTML and returning only essential markdown content via the --md flag, defuddle removes approximately 60-80% of typical web page boilerplate including scripts, stylesheets, and navigation markup. This reduction in token count allows LLMs to process web content faster and with greater context window availability for reasoning tasks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →