How to Convert Images to Draw.io Diagrams Using AI in Next AI Draw.io
Next AI Draw.io converts uploaded images into editable Draw.io diagrams by encoding them as data URLs, processing them through a Next.js API route with an image-capable LLM, and rendering the returned XML through an embedded Draw.io editor.
Next AI Draw.io is an open-source Next.js application that bridges the gap between static images and structured diagrams. By leveraging large language models with vision capabilities, the tool analyzes uploaded PNG or JPEG files and generates corresponding Draw.io XML that you can edit immediately. This guide explains the complete technical implementation based on the actual source code in the DayuanJiang/next-ai-draw-io repository.
Understanding the Architecture
The image-to-diagram conversion follows a six-stage asynchronous pipeline. When a user uploads an image through components/url-input-dialog.tsx, the file converts to a base64 data URL and flows through the Chat API (app/api/chat/route.ts) to an AI provider configured in lib/ai-providers.ts. The system prompt (lib/system-prompts.ts) instructs the model to output valid Draw.io XML, which lib/chat-helpers.ts extracts and wraps for client rendering. Before display, hooks/use-validate-diagram.ts triggers a secondary validation check via app/api/validate-diagram/route.ts to ensure diagram quality, after which hooks/use-diagram-tool-handlers.ts renders the final output in the Draw.io embed.
Step-by-Step Implementation
Step 1: Image Upload and Data URL Conversion
The process begins in components/url-input-dialog.tsx, where users select or drag-and-drop image files. The component reads the file and converts it to a data URL format (data:image/png;base64,…) that can be transmitted as JSON. This encoded string travels to the server via the Chat API endpoint, preserving the image data within the HTTP request body.
Step 2: Server-Side Request Handling
In app/api/chat/route.ts, the POST handler examines the incoming request and extracts image parts from the last user message. The code verifies whether the selected AI model supports the image_generation capability before attaching the image payload. This validation ensures that only vision-capable models receive image data, preventing errors from incompatible providers.
Step 3: AI Provider Integration
The lib/ai-providers.ts module abstracts calls to specific AI SDKs (such as Claude or GPT-4-Vision). When an image is present, the provider invokes the model with a specialized system prompt from lib/system-prompts.ts that explicitly instructs: "Create a diagram that matches the uploaded image and output the result as Draw.io XML." This prompt engineering ensures the model returns structured XML rather than conversational text.
Step 4: Response Parsing and XML Extraction
After the AI generates a response, lib/chat-helpers.ts parses the content to extract the Draw.io XML string. The helper wraps valid XML in a data URL (data:image/svg+xml;base64,…) so the client can render it directly. If parsing fails or validation errors occur, the utility catches exceptions and returns error messages to the UI for user feedback.
Step 5: Diagram Validation Pipeline
Before displaying the diagram, the application runs a quality check through hooks/use-validate-diagram.ts. This hook sends the generated diagram image to app/api/validate-diagram/route.ts, which runs a second LLM configured via lib/validation-schema.ts. The validator checks for visual inconsistencies, structural errors, or missing elements, returning a verdict that determines whether the diagram passes to the rendering stage.
Step 6: Rendering the Final Diagram
Once validation passes, hooks/use-diagram-tool-handlers.ts feeds the Draw.io XML into the Draw.io embed component. The hook manages the integration between the React frontend and the Draw.io editor instance, allowing users to edit, export, or save the AI-generated diagram using native Draw.io controls.
Implementation Examples
Below are practical code snippets showing the critical integration points. Replace YOUR_MODEL_ID with an image-capable model identifier such as gpt-4-vision.
Upload an image from the client:
// Upload component (simplified)
<Button
onClick={async () => {
const file = await pickFile(); // PNG/JPEG selected by user
const base64 = await fileToBase64(file); // Convert to data URL
await fetch('/api/chat', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ image: base64, prompt: 'Create a diagram' })
});
}}
>
Upload Image
</Button>
Process the image on the server:
// Server-side chat handler (core logic)
export async function POST(req: Request) {
const { image, prompt } = await req.json();
// Forward to AI provider (image-enabled model)
const response = await aiProvider.chat({
model: 'YOUR_MODEL_ID',
messages: [
{ role: 'system', content: systemPrompt },
{ role: 'user', content: [{ type: 'image', image }, { type: 'text', text: prompt }] }
]
});
// Extract Draw.io XML from the model's answer
const diagramXml = extractDrawIoXml(response);
return new Response(JSON.stringify({ diagramXml }));
}
Render the generated diagram:
// Rendering the returned diagram
const { diagramXml } = await fetch('/api/chat', …).then(r => r.json());
// Embed Draw.io editor with the generated XML
<DrawIoEditor xml={diagramXml} />
Critical Files and Their Roles
components/url-input-dialog.tsx– Handles image selection and base64 encoding for transmission.app/api/chat/route.ts– Receives image data, validates model capabilities, and forwards to AI providers.lib/ai-providers.ts– Abstracts provider-specific SDK calls for image-capable models.lib/system-prompts.ts– Contains prompts that instruct the AI to output valid Draw.io XML.lib/chat-helpers.ts– Parses AI responses and converts XML to renderable data URLs.hooks/use-validate-diagram.ts– Frontend hook orchestrating diagram quality validation.app/api/validate-diagram/route.ts– Secondary validation endpoint using a separate LLM check.hooks/use-diagram-tool-handlers.ts– Integrates generated XML with the Draw.io editor instance.
Summary
- Image-to-diagram conversion in Next AI Draw.io relies on data URL encoding and vision-capable LLMs to transform static images into structured Draw.io XML.
- The Chat API (
app/api/chat/route.ts) serves as the central hub, routing images to AI providers and handling model capability checks. - System prompts explicitly guide the AI to generate diagram-compatible XML rather than conversational responses.
- Validation pipeline uses a secondary LLM to check diagram quality before rendering, reducing error rates in generated outputs.
- Draw.io integration occurs through specialized React hooks that embed the generated XML directly into the editor interface.
Frequently Asked Questions
What image formats does Next AI Draw.io support?
The application accepts standard web image formats including PNG and JPEG. The components/url-input-dialog.tsx component reads these files as blobs and converts them to base64 data URLs, which the system processes regardless of specific file extension as long as the browser can decode the image.
Which AI models can convert images to diagrams?
You must use models with vision capabilities that support the image_generation flag, such as GPT-4-Vision, Claude 3 Opus, or other multimodal endpoints. The application checks for this capability in app/api/chat/route.ts before sending image payloads, and the provider configuration in lib/ai-providers.ts handles the specific SDK implementation for each supported model.
Why does the application validate diagrams before displaying them?
The validation step in hooks/use-validate-diagram.ts and app/api/validate-diagram/route.ts runs a secondary LLM review to catch structural errors, visual inconsistencies, or malformed XML that the initial generation might produce. This quality gate ensures users receive functional, editable diagrams rather than broken or incomplete Draw.io XML.
Can I edit the generated diagrams after conversion?
Yes. Once the AI generates and validates the Draw.io XML, hooks/use-diagram-tool-handlers.ts renders it within the standard Draw.io embed interface. You can modify shapes, add connections, change colors, and export the diagram just like any native Draw.io creation, as the output uses standard Draw.io XML format compatible with the editor's full feature set.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →