How the AI Email Extraction Feature Utilizes Cloudflare Workers AI in Temporary Email Services
The AI email extraction feature invokes Cloudflare Workers AI directly from the worker runtime to parse incoming messages, extract verification codes and authentication links using a structured JSON schema, and inject the results into notification pipelines.
The cloudflare_temp_email project leverages Cloudflare's serverless AI infrastructure to transform raw email content into actionable data. By integrating Workers AI directly into the email processing pipeline at worker/src/email/index.ts, the service automatically identifies critical information like OTP codes and sign-in links before triggering downstream notifications.
Enabling and Configuring the AI Extraction Pipeline
The extraction workflow begins with environment-level controls and administrative allow-list configurations that determine which incoming messages qualify for AI processing.
Feature Toggle and Environment Variables
The pipeline checks the ENABLE_AI_EMAIL_EXTRACT environment variable before executing any AI logic. In worker/src/email/ai_extract.ts, the extractEmailInfo function returns early if this flag is disabled, preventing unnecessary API calls:
// From worker/src/email/ai_extract.ts
if (!getBooleanValue(env.ENABLE_AI_EMAIL_EXTRACT)) {
return null;
}
The function is invoked from worker/src/email/index.ts after the email has been saved and forwarded but before notifications are dispatched, ensuring extraction results are available for alerting systems.
Allow-List Configuration via Admin API
When enabled, the system loads an allow-list defined in admin_api/ai_extract_settings.ts using the getJsonSetting utility. The recipient address is compared against configured patterns at lines 53-73 of worker/src/email/ai_extract.ts, and only matched addresses trigger the AI call. This prevents token consumption on irrelevant traffic.
Configure the allow-list via the admin API:
# GET current settings
GET /admin_api/ai_extract_settings
# Update settings
POST /admin_api/ai_extract_settings
{
"enableAllowList": true,
"allowList": ["*.example.com", "temp@mydomain.com"]
}
Preparing Email Content for Workers AI
Raw email content requires preprocessing to meet Workers AI input requirements and token limits before the model invocation occurs.
Text Extraction and HTML Conversion
The system first parses the email using commonParseMail. For plain-text emails, parsedEmail.text passes directly to the model. HTML messages undergo processing via htmlToTextForAi (defined in worker/src/email/ai_extract.ts at lines 200-214), which strips scripts and styles, decodes HTML entities, and preserves link text to maintain semantic value while removing markup.
Token Limit Management
To stay within Workers AI context windows, content exceeding 4 KB is truncated at lines 302-305 of worker/src/email/ai_extract.ts:
if (content.length > 4000) {
content = content.substring(0, 4000);
}
This truncation ensures the model receives complete context without hitting limits, particularly important when processing lengthy HTML newsletters or verbose notification emails.
Invoking Cloudflare Workers AI for Structured Extraction
The core integration occurs when the worker runtime calls the AI binding with a carefully crafted prompt designed to produce deterministic, machine-readable output.
Model Selection and System Prompt
When the env.AI binding is present, the code invokes env.AI.run with the model specified in AI_EXTRACT_MODEL (defaulting to @cf/meta/llama-3.1-8b-instruct-fast). The PROMPT constant at lines 17-88 of worker/src/email/ai_extract.ts instructs the model to:
- Analyze the email subject and body context
- Extract the highest-priority item: authentication code, authentication link, service link, subscription link, or other link
- Return strictly formatted JSON
JSON Schema Enforcement
The invocation uses response_format set to a JSON schema that guarantees parsable output:
// From worker/src/email/ai_extract.ts lines 109-125
const answer = await env.AI.run(model, {
messages: [
{ role: 'system', content: PROMPT },
{ role: 'user', content: content }
],
response_format: {
type: 'json',
schema: extractJsonSchema
}
});
This structured approach eliminates parsing ambiguity and ensures the downstream ExtractResult type validation succeeds.
Handling the AI Response
The returned JSON is parsed into an ExtractResult interface. If the result contains meaningful data (type !== 'none' && result.result), the system persists it to the raw_mails table under the metadata.ai_extract field (lines 110-113), making it queryable for future reference and analytics.
Fallback Mechanisms and Downstream Integration
The architecture includes resilience patterns for environments without AI bindings and integrates results into notification workflows.
Regex-Based Fallback Extraction
If the Workers AI binding is unavailable (such as in self-hosted deployments without Cloudflare), the code falls back to a built-in regex extractor at lines 88-96 of worker/src/email/ai_extract.ts. The extractCode function uses pattern matching to pull verification codes, ensuring the feature provides value even without AI infrastructure.
Telegram and Webhook Notifications
The aiExtractResult flows into two primary notification channels defined in worker/src/email/index.ts at lines 29-33 and 38-42:
- Telegram: Passed to
sendMailToTelegramto include extracted codes or links in chat notifications - Webhooks: Forwarded to
triggerWebhookallowing external services to receive structured extraction data via HTTP callbacks
This dual-path integration ensures extracted intelligence reaches end-users and automation systems immediately upon processing.
Summary
- Environment Control: The feature activates via
ENABLE_AI_EMAIL_EXTRACTand uses an admin-managed allow-list to filter eligible recipients. - Content Preparation: HTML conversion via
htmlToTextForAiand 4 KB truncation prepare emails for the AI model. - Workers AI Integration: The system calls
env.AI.runwith a structured JSON schema using the@cf/meta/llama-3.1-8b-instruct-fastmodel by default. - Resilience: Regex fallback extraction ensures functionality when AI bindings are absent.
- Downstream Usage: Results propagate to Telegram notifications and webhook dispatchers for immediate action.
Frequently Asked Questions
How do I enable the AI email extraction feature in my deployment?
Set the environment variable ENABLE_AI_EMAIL_EXTRACT to true in your wrangler.toml or Cloudflare Workers environment. Optionally configure AI_EXTRACT_MODEL to override the default Llama 3.1 model.
What happens if Cloudflare Workers AI is not available in my environment?
The system automatically falls back to a regex-based extractor (extractCode) that scans for verification codes using pattern matching. This ensures basic extraction capabilities persist even without AI infrastructure.
How does the system handle large HTML emails?
Content longer than 4,000 characters is truncated before being sent to Workers AI. HTML is converted to clean text via htmlToTextForAi, which removes scripts and styles while preserving link text and semantic content.
Where are the extraction results stored and how can I access them?
Successful extractions are stored in the raw_mails table under the metadata.ai_extract column as JSON. They are also passed to Telegram notifications and webhook endpoints if those integrations are configured.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →