# How the Web Search Enhancement Feature Uses DuckDuckGo Integration in nGPT

> Discover how nGPT integrates DuckDuckGo search results without an API key. Learn about its parsing, content extraction, and prompt formatting for LLMs.

- Repository: [nazDridoy/ngpt](https://github.com/nazdridoy/ngpt)
- Tags: deep-dive
- Published: 2026-03-07

---

**The web search enhancement feature in nGPT queries DuckDuckGo's public HTML endpoint without requiring an API key, parses results using BeautifulSoup to extract titles and decode redirect URLs, fetches full article content through a hybrid extraction pipeline, and formats everything into an LLM-ready prompt block with optional citation instructions.**

The `nazdridoy/ngpt` repository implements a sophisticated web search enhancement feature that augments LLM prompts with real-time information from DuckDuckGo. This integration operates entirely through the public HTML interface—no API keys required—making it accessible for any nGPT deployment. Understanding how this DuckDuckGo integration works reveals a multi-stage pipeline of search querying, HTML parsing, content extraction, and prompt formatting.

## Core Architecture of the DuckDuckGo Integration

The entire workflow lives in [`ngpt/utils/web_search.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/utils/web_search.py) and centers on three primary operations: issuing search queries, parsing HTML responses, and extracting full article content.

### Querying the DuckDuckGo HTML Endpoint

The `perform_web_search()` function constructs a direct HTTP request to `https://html.duckduckgo.com/html/`. Unlike API-based integrations, this approach uses the public HTML interface, sending a custom **User-Agent** and **Accept** header to mimic a real browser:

```python
url = f"https://html.duckduckgo.com/html/?q={encoded_query}"
response = requests.get(url, headers=headers, timeout=10)

```

This implementation appears in [`ngpt/utils/web_search.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/utils/web_search.py) at lines 40–63. The function respects the `max_results` parameter, returning up to the specified number of search results.

### Parsing Search Results and Decoding Redirects

Once the HTML response arrives, the code uses **BeautifulSoup** with the `html.parser` backend to extract structured data. Each result block uses the CSS selector `.result`, from which the code extracts the title, snippet, and URL.

A critical challenge is that DuckDuckGo wraps outbound links in a redirect mechanism (`/l/?uddg=…`). The parser handles this by extracting the `uddg` query parameter to reveal the actual destination URL:

```python
href = title_elem.find('a')['href']
parsed_url = urlparse(href)
actual_url = parse_qs(parsed_url.query).get('uddg', [None])[0]

```

This URL decoding logic appears in [`ngpt/utils/web_search.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/utils/web_search.py) at lines 73–86. The function then returns a list of dictionaries containing `title`, `href`, and `body` (the snippet) for each result.

## Full Article Extraction Pipeline

Retrieving search result metadata is only the first step. The `extract_article_content()` function performs a second HTTP request to fetch the actual webpage and extract readable content using a hybrid extraction strategy.

### The Hybrid Content Extraction Strategy

Located in [`ngpt/utils/web_search.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/utils/web_search.py) at lines 101–210, this function implements a multi-layered approach:

1. **Sanitization**: Strips scripts, styles, comments, hidden elements, and iframes to remove non-content markup.
2. **Site-Specific Selectors**: Applies targeted CSS selectors for known domains like Wikipedia and major news sites to extract primary content areas directly.
3. **Structured Data Fallback**: Parses JSON-LD schema markup and meta description tags when semantic selectors fail.
4. **Heuristic Scoring**: Evaluates remaining semantic blocks (`article`, `main`, `section`, `div`) based on text density and link-to-text ratios to identify the primary content container.

The function accepts a `max_chars` parameter to limit the returned content length, ensuring the final prompt does not exceed context window limits.

### Structured Data Assembly

The `get_web_search_results()` function orchestrates the search and extraction phases. It calls `perform_web_search()` to retrieve result metadata, then iterates through each result to call `extract_article_content()`. It attaches a timestamp and constructs a final data structure where each result contains `title`, `url`, `snippet`, and `content` (the full extracted text when available).

This assembly logic appears in [`ngpt/utils/web_search.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/utils/web_search.py) at lines 92–108.

## Formatting Results for LLM Consumption

Raw extracted content requires structured formatting before injection into LLM prompts. The module provides specialized formatting functions to create citation-ready context blocks.

### Prompt Injection and Citation Control

The `format_web_search_results_for_prompt()` function transforms the structured search data into a human-readable block suitable for LLM context windows. This block includes:

- A header displaying the original search query and timestamp
- Numbered result sections containing title, URL, and extracted content
- Explicit citation instructions guiding the LLM to reference sources

Located in [`ngpt/utils/web_search.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/utils/web_search.py) at lines 43–78, this function ensures the model receives clear, structured context with proper attribution metadata.

The `enhance_prompt_with_web_search()` function serves as the public API used by CLI and UI layers. It accepts the original user prompt, orchestrates the search and formatting pipeline, and returns the enhanced prompt with web context prepended. This function supports a `disable_citations` parameter (lines 85–127) that short-circuits the standard formatting to produce a minimal context block without citation instructions—useful for code-generation modes where source attribution is unnecessary.

## Practical Usage Examples

### Directly Enhance a Prompt

```python
from ngpt.utils.web_search import enhance_prompt_with_web_search

user_prompt = "Explain the current state of quantum computing research."
enhanced = enhance_prompt_with_web_search(user_prompt, max_results=3)

print(enhanced)   # The printed text now starts with a “[Web Search Results …]” block

```

Internally this calls `perform_web_search()` → `extract_article_content()` → `get_web_search_results()` → `format_web_search_results_for_prompt()`.

### Use Low-Level Helpers

```python
from ngpt.utils.web_search import perform_web_search, extract_article_content

# Get raw DuckDuckGo results

raw = perform_web_search("best Python testing frameworks", max_results=4)

# Pull full article text for each URL

for item in raw:
    article = extract_article_content(item["href"], max_chars=2000)
    print(f"TITLE: {item['title']}\nURL: {item['href']}\nCONTENT:\n{article[:500]}\n{'-'*80}")

```

This snippet showcases the two-step approach without the LLM-specific formatting.

### Disable Citations for Code Generation

```python
enhanced_no_cite = enhance_prompt_with_web_search(
    "Write a Bash script that backs up /etc to /backup.", 
    max_results=2,
    disable_citations=True
)

print(enhanced_no_cite)   # Same web-search block, but no citation policy text

```

The `disable_citations` flag short-circuits `format_web_search_results_for_prompt()` and builds a minimal block (see lines 105–120 in [`ngpt/utils/web_search.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/utils/web_search.py)).

## Summary

- **No API key required**: The integration uses DuckDuckGo's public HTML endpoint (`https://html.duckduckgo.com/html/`) with browser-mimicking headers.
- **Two-stage extraction**: First retrieves search metadata via `perform_web_search()`, then fetches full articles through `extract_article_content()` using a hybrid sanitization and heuristic scoring pipeline.
- **Redirect decoding**: Parses DuckDuckGo's `/l/?uddg=…` wrapper URLs to extract actual destination addresses using `urllib.parse`.
- **Flexible formatting**: `format_web_search_results_for_prompt()` creates citation-ready context blocks, while `enhance_prompt_with_web_search()` provides a simple API with optional citation disabling for code-generation tasks.
- **Centralized implementation**: All functionality resides in [`ngpt/utils/web_search.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/utils/web_search.py), consumed by CLI and UI layers in [`ngpt/cli/main.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/cli/main.py) and [`ngpt/ui/interactive_ui.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/ui/interactive_ui.py).

## Frequently Asked Questions

### Does nGPT require a DuckDuckGo API key to use the web search enhancement feature?

No. The implementation in [`ngpt/utils/web_search.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/utils/web_search.py) uses the public HTML interface at `https://html.duckduckgo.com/html/`. The `perform_web_search()` function sends standard HTTP GET requests with custom User-Agent headers to mimic a browser, eliminating the need for API keys or authentication tokens.

### How does nGPT extract actual URLs from DuckDuckGo's search results?

DuckDuckGo wraps outbound links in a redirect mechanism using the path `/l/?uddg=…`. The parsing logic in `perform_web_search()` (lines 73–86) extracts the `href` attribute, parses it with `urllib.parse.urlparse()`, and decodes the `uddg` query parameter to reveal the final destination URL.

### What methods does nGPT use to extract readable content from web pages?

The `extract_article_content()` function implements a hybrid extraction strategy. It first sanitizes HTML by removing scripts, styles, and hidden elements. Then it applies site-specific CSS selectors for domains like Wikipedia, falls back to JSON-LD structured data or meta descriptions, and finally uses heuristic scoring of semantic blocks (`article`, `main`, `section`) based on text density to identify primary content.

### Can I disable citation instructions when using web search for code generation tasks?

Yes. The `enhance_prompt_with_web_search()` function accepts a `disable_citations` parameter. When set to `True`, the function short-circuits the standard formatting logic and builds a minimal context block containing only the search results without citation policy text, which is ideal for code-generation modes where source attribution is unnecessary.