How the AI Email Extraction Feature Utilizes Cloudflare Workers AI in Temporary Email Services

The AI email extraction feature invokes Cloudflare Workers AI directly from the worker runtime to parse incoming messages, extract verification codes and authentication links using a structured JSON schema, and inject the results into notification pipelines.

The cloudflare_temp_email project leverages Cloudflare's serverless AI infrastructure to transform raw email content into actionable data. By integrating Workers AI directly into the email processing pipeline at worker/src/email/index.ts, the service automatically identifies critical information like OTP codes and sign-in links before triggering downstream notifications.

Enabling and Configuring the AI Extraction Pipeline

The extraction workflow begins with environment-level controls and administrative allow-list configurations that determine which incoming messages qualify for AI processing.

Feature Toggle and Environment Variables

The pipeline checks the ENABLE_AI_EMAIL_EXTRACT environment variable before executing any AI logic. In worker/src/email/ai_extract.ts, the extractEmailInfo function returns early if this flag is disabled, preventing unnecessary API calls:

// From worker/src/email/ai_extract.ts
if (!getBooleanValue(env.ENABLE_AI_EMAIL_EXTRACT)) {
    return null;
}

The function is invoked from worker/src/email/index.ts after the email has been saved and forwarded but before notifications are dispatched, ensuring extraction results are available for alerting systems.

Allow-List Configuration via Admin API

When enabled, the system loads an allow-list defined in admin_api/ai_extract_settings.ts using the getJsonSetting utility. The recipient address is compared against configured patterns at lines 53-73 of worker/src/email/ai_extract.ts, and only matched addresses trigger the AI call. This prevents token consumption on irrelevant traffic.

Configure the allow-list via the admin API:


# GET current settings

GET /admin_api/ai_extract_settings

# Update settings

POST /admin_api/ai_extract_settings
{
  "enableAllowList": true,
  "allowList": ["*.example.com", "temp@mydomain.com"]
}

Preparing Email Content for Workers AI

Raw email content requires preprocessing to meet Workers AI input requirements and token limits before the model invocation occurs.

Text Extraction and HTML Conversion

The system first parses the email using commonParseMail. For plain-text emails, parsedEmail.text passes directly to the model. HTML messages undergo processing via htmlToTextForAi (defined in worker/src/email/ai_extract.ts at lines 200-214), which strips scripts and styles, decodes HTML entities, and preserves link text to maintain semantic value while removing markup.

Token Limit Management

To stay within Workers AI context windows, content exceeding 4 KB is truncated at lines 302-305 of worker/src/email/ai_extract.ts:

if (content.length > 4000) {
    content = content.substring(0, 4000);
}

This truncation ensures the model receives complete context without hitting limits, particularly important when processing lengthy HTML newsletters or verbose notification emails.

Invoking Cloudflare Workers AI for Structured Extraction

The core integration occurs when the worker runtime calls the AI binding with a carefully crafted prompt designed to produce deterministic, machine-readable output.

Model Selection and System Prompt

When the env.AI binding is present, the code invokes env.AI.run with the model specified in AI_EXTRACT_MODEL (defaulting to @cf/meta/llama-3.1-8b-instruct-fast). The PROMPT constant at lines 17-88 of worker/src/email/ai_extract.ts instructs the model to:

  • Analyze the email subject and body context
  • Extract the highest-priority item: authentication code, authentication link, service link, subscription link, or other link
  • Return strictly formatted JSON

JSON Schema Enforcement

The invocation uses response_format set to a JSON schema that guarantees parsable output:

// From worker/src/email/ai_extract.ts lines 109-125
const answer = await env.AI.run(model, {
    messages: [
        { role: 'system', content: PROMPT },
        { role: 'user', content: content }
    ],
    response_format: {
        type: 'json',
        schema: extractJsonSchema
    }
});

This structured approach eliminates parsing ambiguity and ensures the downstream ExtractResult type validation succeeds.

Handling the AI Response

The returned JSON is parsed into an ExtractResult interface. If the result contains meaningful data (type !== 'none' && result.result), the system persists it to the raw_mails table under the metadata.ai_extract field (lines 110-113), making it queryable for future reference and analytics.

Fallback Mechanisms and Downstream Integration

The architecture includes resilience patterns for environments without AI bindings and integrates results into notification workflows.

Regex-Based Fallback Extraction

If the Workers AI binding is unavailable (such as in self-hosted deployments without Cloudflare), the code falls back to a built-in regex extractor at lines 88-96 of worker/src/email/ai_extract.ts. The extractCode function uses pattern matching to pull verification codes, ensuring the feature provides value even without AI infrastructure.

Telegram and Webhook Notifications

The aiExtractResult flows into two primary notification channels defined in worker/src/email/index.ts at lines 29-33 and 38-42:

  • Telegram: Passed to sendMailToTelegram to include extracted codes or links in chat notifications
  • Webhooks: Forwarded to triggerWebhook allowing external services to receive structured extraction data via HTTP callbacks

This dual-path integration ensures extracted intelligence reaches end-users and automation systems immediately upon processing.

Summary

  • Environment Control: The feature activates via ENABLE_AI_EMAIL_EXTRACT and uses an admin-managed allow-list to filter eligible recipients.
  • Content Preparation: HTML conversion via htmlToTextForAi and 4 KB truncation prepare emails for the AI model.
  • Workers AI Integration: The system calls env.AI.run with a structured JSON schema using the @cf/meta/llama-3.1-8b-instruct-fast model by default.
  • Resilience: Regex fallback extraction ensures functionality when AI bindings are absent.
  • Downstream Usage: Results propagate to Telegram notifications and webhook dispatchers for immediate action.

Frequently Asked Questions

How do I enable the AI email extraction feature in my deployment?

Set the environment variable ENABLE_AI_EMAIL_EXTRACT to true in your wrangler.toml or Cloudflare Workers environment. Optionally configure AI_EXTRACT_MODEL to override the default Llama 3.1 model.

What happens if Cloudflare Workers AI is not available in my environment?

The system automatically falls back to a regex-based extractor (extractCode) that scans for verification codes using pattern matching. This ensures basic extraction capabilities persist even without AI infrastructure.

How does the system handle large HTML emails?

Content longer than 4,000 characters is truncated before being sent to Workers AI. HTML is converted to clean text via htmlToTextForAi, which removes scripts and styles while preserving link text and semantic content.

Where are the extraction results stored and how can I access them?

Successful extractions are stored in the raw_mails table under the metadata.ai_extract column as JSON. They are also passed to Telegram notifications and webhook endpoints if those integrations are configured.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →