How Bella OpenAPI Simulates Function Calling for LLMs Without Native Support
Bella OpenAPI bridges the capability gap for non-native models by intercepting function-calling requests, rewriting them as Python-style code generation prompts, and parsing the LLM's text output back into structured tool-call objects, enabling full compatibility with Claude, Gemini, and other models.
Bella OpenAPI serves as a unified gateway for heterogeneous Large Language Models, but providers like Anthropic's Claude or Google's Gemini may lack native OpenAI-style function_call support. For these scenarios, the lianjiatech/bella-openapi repository implements a transparent simulation layer that converts standard function calling protocols into prompt-based code generation tasks, then reconstructs the responses into the expected tool-call format without requiring changes to downstream services.
Activating Function Calling Simulation
Detection via HTTP Headers and Properties
The simulation workflow begins when the client explicitly requests simulation mode. In ChatController.java, the system checks for the X-BELLA-FUNCTION-SIMULATE header or the functionCallSimulate flag within CompletionProperty to determine if the target model requires prompt-based simulation.
// From ChatController.java
String functionCallSimulate = requestInfo.get("X-BELLA-FUNCTION-SIMULATE");
property.setFunctionCallSimulate("true".equals(functionCallSimulate) || property.isFunctionCallSimulate());
When this flag is set and the request contains functions or tools definitions, the ToolCallSimulator intercepts the request before it reaches the underlying model adaptor.
Request Rewriting with Pebble Templates
The core transformation occurs in SimulationHepler.java, where the rewrite(request) method constructs a specialized prompt describing the available functions as Python signatures. This prompt is rendered from the Pebble template function_call_template.pebble and injected as a single user message that instructs the LLM to generate executable Python code.
// From SimulationHepler.java
String prompt = Renders.render(
"com/ke/bella/openapi/simulation/function_call_template.pebble",
env);
Message msg = Message.builder().role("user").content(prompt).build();
The rewritten request asks the model to emit a Python function call matching one of the supplied signatures, effectively turning any text-completion model into a function-calling agent.
Parsing and Converting Responses
Python Code Parsing
After the LLM generates the raw text response (placed in the content field), SimulationHepler.parse() executes a lightweight PythonFuncCallParser to analyze the output. If the parser detects a valid function call, it constructs a CompletionResponse.Choice containing a structured tool-call object; otherwise, the text passes through as a standard assistant message.
// From SimulationHepler.java
Choice choice = SimulationHepler.parse(resp.reasoning(), resp.content());
// choice now carries Message.ToolCall with FunctionCall{name, arguments}
Adapter Integration
ToolCallSimulator swaps the first choice of the original response with the parsed tool-call choice while preserving finish reasons and metadata. Converters like VertexConverter and ResponsesApiConverter then process these internal Message.ToolCall objects exactly as they would native function-call responses, ensuring that quota handling, logging, and other downstream services remain unaffected.
// Logic flow in ToolCallSimulator
processData.setFunctionCallSimulate(true); // Flags downstream adaptors
// Swaps response choices to inject parsed tool calls
Streaming Support for Simulated Calls
For streaming endpoints, the simulator applies the same rewriting logic before delegating to the underlying adaptor. The processData.setFunctionCallSimulate(true) flag ensures that streaming handlers recognize the simulation context and properly format the incremental tool-call chunks as they arrive from providers that do not natively support function calling.
Summary
- Bella OpenAPI enables function calling on non-native models through a prompt-based simulation layer that requires no client-side changes to downstream services.
- Activation occurs via the
X-BELLA-FUNCTION-SIMULATEheader orCompletionPropertyflags, detected inChatController.java. - Transformation happens in
SimulationHepler.java, which renders Python-style function signatures using thefunction_call_template.pebbletemplate. - Parsing relies on
PythonFuncCallParserto extract structured tool calls from raw Python code generated by the LLM. - Compatibility is maintained through
ToolCallSimulator, which integrates withVertexConverterandResponsesApiConverterto produce standard OpenAI-compatible responses. - Streaming is fully supported by setting simulation flags in the process data before delegation to the underlying adaptor.
Frequently Asked Questions
Which LLM providers benefit from this simulation mechanism?
Models such as Anthropic Claude, Google Gemini, and other text-completion engines that lack native OpenAI-style function_call or tools parameters utilize this layer. The simulation allows these models to participate fully in Bella OpenAPI's unified gateway features without requiring provider-specific client code or API modifications.
How does the prompt engineering work for function simulation?
The system uses the Pebble templating engine to render function_call_template.pebble, which transforms the JSON schema of available functions into Python function signatures and usage instructions. This template is dynamically populated with the agent context and prior tool-call history, then sent to the LLM as a standard user message requesting Python code output that matches the specified signatures.
Does function calling simulation impact response latency?
Yes, simulation introduces minimal overhead due to the additional parsing step. The PythonFuncCallParser must analyze the raw text output, and the request rewriting adds a preprocessing phase. However, because the underlying LLM generates code in a single pass rather than through iterative API calls, the total latency remains comparable to standard completion requests while adding full tool-use capabilities to previously unsupported models.
Can the simulator handle multiple function calls in a single response?
The SimulationHepler.parse() method and PythonFuncCallParser are designed to detect single function calls from the generated Python code. Complex multi-call scenarios depend on the specific implementation details in ToolCallSimulator, which manages the Choice object reconstruction. Currently, the architecture focuses on accurately parsing the first valid function call detected in the LLM's output and treating subsequent text as standard content.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →