How the y-gui Tool Parser Extracts and Executes Tools from AI Responses

The y-gui tool parser detects XML tool blocks in LLM responses using lightweight string searches, isolates them from conversational text, and extracts server names, tool names, and arguments via regex and JSON5 parsing before routing execution to the MCP manager.

The y-gui open-source project implements a robust pipeline for handling AI-generated tool requests embedded in chat responses. When an LLM returns a message containing a <use_mcp_tool> block, the tool parser systematically extracts the XML payload, separates it from user-facing content, and prepares the parameters for secure execution through the Model Context Protocol (MCP).

The parsing pipeline begins with a cheap, synchronous check to determine if the assistant's message contains a tool request. In backend/src/utils/tool-parser.ts, the containsToolUse() function performs a simple string search for the <use_mcp_tool> opening and closing tags.

export function containsToolUse(content: string): boolean {
  const toolTags = ["use_mcp_tool"];
  for (const tag of toolTags) {
    if (content.includes(`<${tag}>`) && content.includes(`</${tag}>`)) {
      return true;
    }
  }
  return false;
}

This approach avoids expensive parsing operations when no tool invocation is present. If the function returns true, the pipeline proceeds to isolate the XML payload from the conversational text.

Splitting Conversational Text from XML Payloads

Once a tool block is detected, the splitContent() function in backend/src/utils/tool-parser.ts divides the raw message into two distinct parts:

  • plainContent: Human-readable text displayed to the user
  • toolContent: The raw XML block containing the tool invocation
export function splitContent(content: string): [string, string | null] {
  const toolTags = ["use_mcp_tool"];
  // locate the first opening tag …
  // locate the matching closing tag …
  // return [plainContent, toolContent];
}

This separation ensures that users see only the conversational explanation while the system extracts the structured tool data hidden within the XML tags.

Extracting Server, Tool Name, and Arguments

The core extraction logic resides in extractMcpToolUse(), which uses targeted regular expressions to parse the isolated XML block. The function specifically looks for <server_name>, <tool_name>, and <arguments> tags, using the JSON5 library to parse the arguments payload with tolerance for trailing commas and JavaScript-style comments.

export function extractMcpToolUse(content: string): [string, string, any] | null {
  const mcpToolMatch = content.match(/<use_mcp_tool>(.*?)<\/use_mcp_tool>/s);
  if (!mcpToolMatch) return null;
  const toolContent = mcpToolMatch[1];

  const serverMatch = toolContent.match(/<server_name>(.*?)<\/server_name>/);
  const toolMatch   = toolContent.match(/<tool_name>(.*?)<\/tool_name>/);
  const argsMatch   = toolContent.match(/<arguments>\s*(\{.*?\})\s*<\/arguments>/s);

  if (!serverMatch || !toolMatch || !argsMatch) return null;

  const args = JSON5.parse(argsMatch[1]);
  return [serverMatch[1].trim(), toolMatch[1].trim(), args];
}

This function returns a tuple containing the server name, tool name, and parsed arguments object, or null if any required field is missing.

Integrating the Parser into the Chat Flow

The ChatService.processAssistantMessage method in backend/src/serivce/chat.ts orchestrates the complete parsing workflow through five sequential steps:

  1. Detect: Call containsToolUse(content) to check for tool blocks
  2. Split: Use splitContent(content) to separate user text from XML
  3. Extract: Invoke extractMcpToolUse(toolContent!) to parse parameters
  4. Update: Enrich the Message object with server, tool, and arguments properties
  5. Confirm: Stream a confirmationRequest payload to the frontend for user approval
if (!containsToolUse(content)) { /* handle regular message */ }

const [plainContent, toolContent] = splitContent(content);
assistantMessage.content = plainContent.trim() ? plainContent : "I'll execute this operation for you.";

const mcpTool = extractMcpToolUse(toolContent!);
if (mcpTool) {
  const [serverName, toolName, args] = mcpTool;
  assistantMessage.server = serverName;
  assistantMessage.tool = toolName;
  assistantMessage.arguments = args;

  const confirmationRequest = { 
    plainContent, 
    server: serverName, 
    tool: toolName, 
    arguments: args 
  };
  await writer.write(encoder.encode(`data: ${JSON.stringify(confirmationRequest)}\n\n`));
}

The confirmation payload includes the plain text explanation alongside the structured tool data, allowing the frontend UI to display a "Run tool" button that requires explicit user consent before execution proceeds.

Executing Tools After User Confirmation

Following user confirmation via the frontend, the tool execution moves outside the parser's scope. The frontend calls the /api/tool/confirm endpoint exposed in backend/src/api/tool-router.ts, which delegates to the McpManager class in backend/src/mcp/mcp-manager.ts. This manager looks up the specified MCP server by name, invokes the remote tool with the parsed arguments, and returns the result to the client as a new assistant message in the chat history.

Summary

  • Detection: The parser uses lightweight string searches in containsToolUse() to identify <use_mcp_tool> blocks without expensive XML parsing.
  • Separation: The splitContent() function isolates conversational text from XML payloads to maintain clean UI rendering.
  • Extraction: Regex-based parsing in extractMcpToolUse() extracts server names, tool names, and JSON5-parsed arguments from the XML structure.
  • Integration: The chat service orchestrates detection, splitting, and extraction before streaming confirmation payloads to the frontend.
  • Execution: User confirmation triggers the McpManager to invoke the actual remote tool through the Model Context Protocol.

Frequently Asked Questions

How does y-gui handle malformed JSON in tool arguments?

The parser uses JSON5 rather than standard JSON.parse() to process the content within <arguments> tags. JSON5 tolerates trailing commas, single quotes, and JavaScript-style comments, making the system more resilient to LLM-generated formatting errors.

What happens if the LLM response contains multiple tool blocks?

The current implementation in splitContent() locates the first opening <use_mcp_tool> tag and its matching closing tag, processing only the first tool block found. Subsequent tool blocks in the same message would require additional parsing passes or message splitting logic.

Where is the user confirmation step implemented in the codebase?

User confirmation occurs in backend/src/serivce/chat.ts within the processAssistantMessage method, which streams a confirmationRequest JSON payload to the frontend via Server-Sent Events. The frontend displays this as a pending tool execution request, and only upon user approval does the frontend call the /api/tool/confirm endpoint to trigger actual execution via backend/src/mcp/mcp-manager.ts.

Can the tool parser handle nested XML or HTML content within arguments?

The extractMcpToolUse() function uses a non-greedy regex pattern (\{.*?\}) with the s flag to capture JSON content within <arguments> tags. This approach assumes the arguments contain valid JSON or JSON5 objects and does not support arbitrary nested XML structures within the arguments field.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →