How FreeLLMAPI Handles Tool-Call Rescue for Incorrect Emissions
FreeLLMAPI intercepts malformed tool-call emissions from large language models and repairs them through a multi-stage pipeline involving marker detection, payload buffering, and argument reconstruction.
FreeLLMAPI serves as a robust proxy layer for diverse LLM providers that may deviate from standard tool-call syntax. The tool-call rescue system, maintained in the tashfeenahmed/freellmapi repository, automatically detects and repairs incorrect emissions to maintain API compatibility without breaking client-side integrations.
The Tool-Call Rescue Pipeline
The rescue mechanism operates through five coordinated stages when processing streaming responses from model providers.
Detection of Inline Markers
The system monitors incoming token streams for tool-call dialect markers using helper functions defined in server/src/lib/tool-call-rescue.ts. The utilities startsWithDialectMarker, couldBecomeDialectMarker, and containsDialectMarker scan for the <<tool-call>> sequence or OpenAI-style function_call blocks that indicate the start of a tool invocation payload.
Holding Incomplete Payloads
When a marker is detected but the JSON payload remains incomplete, the resolver enters a holding state to buffer subsequent tokens. This logic appears in server/src/routes/responses.ts around line 788 and server/src/routes/proxy.ts around line 2058. The holding mechanism prevents stray prose from contaminating the reconstruction buffer while waiting for the closing brace or delimiter that completes the tool-call object.
Repairing Malformed Arguments
Once fragment boundaries are identified, the repairToolArguments function in server/src/lib/tool-args.ts sanitizes the captured string. The repair logic corrects double-encoded JSON, inserts missing quotes around keys, and normalizes the structure to match the declared function schema. For example, a malformed fragment like {city: Paris transforms into valid {"city":"Paris"} syntax.
Assembly and Event Injection
After successful parsing, the rescued call is assembled into a standard function_call event conforming to the OpenAI API specification. The assembly code in responses.ts near line 1027 and proxy.ts near line 2271 injects this reconstructed event into the streaming response. Downstream consumers receive a well-formed tool-call object despite the original model emission containing syntax errors.
Error Handling and Fallback Triggers
If the rescue process cannot produce valid JSON, the system throws an unparseable inline tool-call dialect from <route> error. As implemented in server/src/lib/error-classify.ts at line 99, this error classification feeds into the fallback logic in server/src/lib/fallback-loop.ts at line 949. When rescue failures accumulate, the fallback loop automatically switches to alternative model providers that may produce more reliable tool-call syntax.
Activation and Scope
The rescue system activates only for requests that explicitly require tool usage. The check for requirements.requireTools in server/src/routes/proxy.ts at line 1602 ensures the overhead of marker detection and payload holding is incurred only when necessary, preventing false-positive parsing in pure chat scenarios.
Practical Implementation Example
The following pattern demonstrates how the rescue mechanism processes a streaming response containing a malformed tool-call emission:
import { rescueInlineToolCalls } from '@/lib/tool-call-rescue';
// Process a raw chunk from the LLM stream
const rescued = rescueInlineToolCalls({
text: rawChunk,
route: 'v1/chat/completions',
requirements: { requireTools: true },
});
if (rescued.calls && rescued.calls.length > 0) {
// Emit standardized function_call events to the client
for (const call of rescued.calls) {
stream.write(`data: ${JSON.stringify({
choices: [{ delta: { function_call: call } }]
})}\n\n`);
}
}
When a client sends a request requiring tool usage:
await fetch('https://api.freellm.ai/v1/chat/completions', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
model: 'gpt-oss-120b',
messages: [{ role: 'user', content: 'Get weather for Paris' }],
tools: [{
type: 'function',
function: {
name: 'get_weather',
parameters: {
type: 'object',
properties: { city: { type: 'string' } }
}
}
}]
})
});
If the model emits a malformed payload like {"name": "get_weather", "arguments": "{city: Paris", the rescue pipeline detects the opening marker, holds the stream until the closing brace, repairs the argument string to {"city":"Paris"}, and injects the corrected function_call event into the response stream.
Summary
- FreeLLMAPI implements tool-call rescue in
server/src/lib/tool-call-rescue.tsto handle malformed LLM emissions without breaking client integrations. - The detection phase uses
containsDialectMarkerand related helpers to identify<<tool-call>>sequences in streaming tokens. - Holding logic in
responses.tsandproxy.tsbuffers incomplete payloads until parsable JSON fragments are received. - The repair phase employs
repairToolArgumentsfromserver/src/lib/tool-args.tsto fix JSON syntax errors and schema mismatches. - Failed rescues trigger fallback mechanisms via
error-classify.tsandfallback-loop.tsto switch to alternative model providers. - Rescue processing is conditional on
requirements.requireToolsto avoid overhead in standard chat requests.
Frequently Asked Questions
What triggers the tool-call rescue mechanism in FreeLLMAPI?
The mechanism triggers when a streaming response contains text that matches tool-call dialect markers while the request has requirements.requireTools set to true. The system scans for <<tool-call>> markers or partial function_call JSON objects, activating the rescue pipeline upon detection of incomplete or malformed syntax.
How does FreeLLMAPI repair malformed tool arguments?
FreeLLMAPI repairs arguments through the repairToolArguments function in server/src/lib/tool-args.ts. This utility fixes common JSON syntax errors including double-encoded strings, unquoted object keys, and truncated values. It normalizes the repaired content against the function's declared schema to ensure downstream validation passes.
What happens when tool-call rescue fails?
When rescue fails to produce valid JSON, the system throws an unparseable inline tool-call dialect error. This error is classified in server/src/lib/error-classify.ts and consumed by the fallback loop in server/src/lib/fallback-loop.ts. The fallback system then routes subsequent requests to alternative model providers that may emit correct tool-call syntax.
Is tool-call rescue active for all requests?
No, tool-call rescue is conditionally activated. The code in server/src/routes/proxy.ts explicitly checks requirements.requireTools before engaging the rescue pipeline. This optimization ensures that standard chat requests without tool requirements bypass the overhead of marker detection and payload reconstruction.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →