How the Web Search Enhancement Feature Uses DuckDuckGo Integration in nGPT
The web search enhancement feature in nGPT queries DuckDuckGo's public HTML endpoint without requiring an API key, parses results using BeautifulSoup to extract titles and decode redirect URLs, fetches full article content through a hybrid extraction pipeline, and formats everything into an LLM-ready prompt block with optional citation instructions.
The nazdridoy/ngpt repository implements a sophisticated web search enhancement feature that augments LLM prompts with real-time information from DuckDuckGo. This integration operates entirely through the public HTML interface—no API keys required—making it accessible for any nGPT deployment. Understanding how this DuckDuckGo integration works reveals a multi-stage pipeline of search querying, HTML parsing, content extraction, and prompt formatting.
Core Architecture of the DuckDuckGo Integration
The entire workflow lives in ngpt/utils/web_search.py and centers on three primary operations: issuing search queries, parsing HTML responses, and extracting full article content.
Querying the DuckDuckGo HTML Endpoint
The perform_web_search() function constructs a direct HTTP request to https://html.duckduckgo.com/html/. Unlike API-based integrations, this approach uses the public HTML interface, sending a custom User-Agent and Accept header to mimic a real browser:
url = f"https://html.duckduckgo.com/html/?q={encoded_query}"
response = requests.get(url, headers=headers, timeout=10)
This implementation appears in ngpt/utils/web_search.py at lines 40–63. The function respects the max_results parameter, returning up to the specified number of search results.
Parsing Search Results and Decoding Redirects
Once the HTML response arrives, the code uses BeautifulSoup with the html.parser backend to extract structured data. Each result block uses the CSS selector .result, from which the code extracts the title, snippet, and URL.
A critical challenge is that DuckDuckGo wraps outbound links in a redirect mechanism (/l/?uddg=…). The parser handles this by extracting the uddg query parameter to reveal the actual destination URL:
href = title_elem.find('a')['href']
parsed_url = urlparse(href)
actual_url = parse_qs(parsed_url.query).get('uddg', [None])[0]
This URL decoding logic appears in ngpt/utils/web_search.py at lines 73–86. The function then returns a list of dictionaries containing title, href, and body (the snippet) for each result.
Full Article Extraction Pipeline
Retrieving search result metadata is only the first step. The extract_article_content() function performs a second HTTP request to fetch the actual webpage and extract readable content using a hybrid extraction strategy.
The Hybrid Content Extraction Strategy
Located in ngpt/utils/web_search.py at lines 101–210, this function implements a multi-layered approach:
- Sanitization: Strips scripts, styles, comments, hidden elements, and iframes to remove non-content markup.
- Site-Specific Selectors: Applies targeted CSS selectors for known domains like Wikipedia and major news sites to extract primary content areas directly.
- Structured Data Fallback: Parses JSON-LD schema markup and meta description tags when semantic selectors fail.
- Heuristic Scoring: Evaluates remaining semantic blocks (
article,main,section,div) based on text density and link-to-text ratios to identify the primary content container.
The function accepts a max_chars parameter to limit the returned content length, ensuring the final prompt does not exceed context window limits.
Structured Data Assembly
The get_web_search_results() function orchestrates the search and extraction phases. It calls perform_web_search() to retrieve result metadata, then iterates through each result to call extract_article_content(). It attaches a timestamp and constructs a final data structure where each result contains title, url, snippet, and content (the full extracted text when available).
This assembly logic appears in ngpt/utils/web_search.py at lines 92–108.
Formatting Results for LLM Consumption
Raw extracted content requires structured formatting before injection into LLM prompts. The module provides specialized formatting functions to create citation-ready context blocks.
Prompt Injection and Citation Control
The format_web_search_results_for_prompt() function transforms the structured search data into a human-readable block suitable for LLM context windows. This block includes:
- A header displaying the original search query and timestamp
- Numbered result sections containing title, URL, and extracted content
- Explicit citation instructions guiding the LLM to reference sources
Located in ngpt/utils/web_search.py at lines 43–78, this function ensures the model receives clear, structured context with proper attribution metadata.
The enhance_prompt_with_web_search() function serves as the public API used by CLI and UI layers. It accepts the original user prompt, orchestrates the search and formatting pipeline, and returns the enhanced prompt with web context prepended. This function supports a disable_citations parameter (lines 85–127) that short-circuits the standard formatting to produce a minimal context block without citation instructions—useful for code-generation modes where source attribution is unnecessary.
Practical Usage Examples
Directly Enhance a Prompt
from ngpt.utils.web_search import enhance_prompt_with_web_search
user_prompt = "Explain the current state of quantum computing research."
enhanced = enhance_prompt_with_web_search(user_prompt, max_results=3)
print(enhanced) # The printed text now starts with a “[Web Search Results …]” block
Internally this calls perform_web_search() → extract_article_content() → get_web_search_results() → format_web_search_results_for_prompt().
Use Low-Level Helpers
from ngpt.utils.web_search import perform_web_search, extract_article_content
# Get raw DuckDuckGo results
raw = perform_web_search("best Python testing frameworks", max_results=4)
# Pull full article text for each URL
for item in raw:
article = extract_article_content(item["href"], max_chars=2000)
print(f"TITLE: {item['title']}\nURL: {item['href']}\nCONTENT:\n{article[:500]}\n{'-'*80}")
This snippet showcases the two-step approach without the LLM-specific formatting.
Disable Citations for Code Generation
enhanced_no_cite = enhance_prompt_with_web_search(
"Write a Bash script that backs up /etc to /backup.",
max_results=2,
disable_citations=True
)
print(enhanced_no_cite) # Same web-search block, but no citation policy text
The disable_citations flag short-circuits format_web_search_results_for_prompt() and builds a minimal block (see lines 105–120 in ngpt/utils/web_search.py).
Summary
- No API key required: The integration uses DuckDuckGo's public HTML endpoint (
https://html.duckduckgo.com/html/) with browser-mimicking headers. - Two-stage extraction: First retrieves search metadata via
perform_web_search(), then fetches full articles throughextract_article_content()using a hybrid sanitization and heuristic scoring pipeline. - Redirect decoding: Parses DuckDuckGo's
/l/?uddg=…wrapper URLs to extract actual destination addresses usingurllib.parse. - Flexible formatting:
format_web_search_results_for_prompt()creates citation-ready context blocks, whileenhance_prompt_with_web_search()provides a simple API with optional citation disabling for code-generation tasks. - Centralized implementation: All functionality resides in
ngpt/utils/web_search.py, consumed by CLI and UI layers inngpt/cli/main.pyandngpt/ui/interactive_ui.py.
Frequently Asked Questions
Does nGPT require a DuckDuckGo API key to use the web search enhancement feature?
No. The implementation in ngpt/utils/web_search.py uses the public HTML interface at https://html.duckduckgo.com/html/. The perform_web_search() function sends standard HTTP GET requests with custom User-Agent headers to mimic a browser, eliminating the need for API keys or authentication tokens.
How does nGPT extract actual URLs from DuckDuckGo's search results?
DuckDuckGo wraps outbound links in a redirect mechanism using the path /l/?uddg=…. The parsing logic in perform_web_search() (lines 73–86) extracts the href attribute, parses it with urllib.parse.urlparse(), and decodes the uddg query parameter to reveal the final destination URL.
What methods does nGPT use to extract readable content from web pages?
The extract_article_content() function implements a hybrid extraction strategy. It first sanitizes HTML by removing scripts, styles, and hidden elements. Then it applies site-specific CSS selectors for domains like Wikipedia, falls back to JSON-LD structured data or meta descriptions, and finally uses heuristic scoring of semantic blocks (article, main, section) based on text density to identify primary content.
Can I disable citation instructions when using web search for code generation tasks?
Yes. The enhance_prompt_with_web_search() function accepts a disable_citations parameter. When set to True, the function short-circuits the standard formatting logic and builds a minimal context block containing only the search results without citation policy text, which is ideal for code-generation modes where source attribution is unnecessary.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →