How Tool Use Works in the Textual Reasoner: Implementation and Best Practices
Tool use in the Textual Reasoner enables supported LLMs to execute graph-analysis functions on Mermaid flowcharts by passing the --tool_use flag, triggering a multi-step pipeline where the model can invoke Python functions like get_number_of_nodes to answer structural queries accurately.
The Textual Reasoner in the junyiye/textflow repository provides an optional tool use capability that allows large language models (LLMs) to perform precise graph analysis on flowchart representations. When enabled, this feature transforms natural language questions about flowchart structure into executable function calls, bridging the gap between textual reasoning and computational graph operations.
How Tool Use Works in the Textual Reasoner
Enabling the Tool Use Flag
In src/reasoner.py, the system exposes a boolean command-line flag --tool_use (lines 42-45) that initiates the tool-calling pathway. When this flag is present, the argument parser stores True in args.tool_use, signaling that the incoming Mermaid representation should be treated as a traversable graph object rather than plain text.
Flag Propagation to ModelWrapper
The flag value propagates through the inference pipeline in src/reasoner.py (lines 98-105). If tool_use is enabled, the ModelWrapper.generate_response method receives both the text prompt and the representation argument (containing the Mermaid code). Without the flag, only the prompt is forwarded, bypassing the tool-use logic entirely.
Detection Logic in ModelWrapper
Inside src/models/model_loader.py, the generate_response method (lines 28-34) automatically detects tool-use eligibility by checking whether a representation was supplied. The presence of a representation implicitly toggles tool_use = True, ensuring that graph data is available when the LLM attempts function calls, even if the explicit flag was omitted.
The LLM Tool-Calling Implementation
For API-based models including gpt-4o and gpt-4o-mini, the system invokes generate_api_response_tool_use from src/models/api_models.py (lines 59-66). This function implements a two-phase protocol: it first sends the prompt without tools, and only if the LLM returns null content does it fall back to the tool-calling pathway.
Tool Definitions and Execution Loop
The available tools are defined in src/models/prompt_utils.py (lines 13-30) via load_tools(), which returns OpenAI function-calling schema specifications for operations like get_number_of_nodes and get_direct_successors. When the LLM returns tool_calls, src/models/api_models.py (lines 86-119) executes load_tools_code() to generate Python code that forwards calls to the flowchart object. Each call runs via exec(), with results captured and appended as tool messages. A second LLM request then processes these results alongside the original prompt to generate the final answer.
When to Enable Tool Use in the Textual Reasoner
Enable the --tool_use flag when your workflow meets specific criteria regarding query complexity, model compatibility, and performance tolerance.
Complex Graph Queries
Activate tool use when downstream questions require structural computation that natural language processing cannot reliably perform, such as counting nodes ("How many nodes are there?"), traversing paths ("What is the shortest path between X and Y?"), or analyzing connectivity. The explicit function calls eliminate ambiguity in numerical or topological answers.
Supported LLM Models
Currently, only gpt-4o and gpt-4o-mini implement the tool-use pathway in src/models/api_models.py. Enabling the flag with unsupported models triggers a silent fallback to standard text generation, as the ModelWrapper detects the unsupported configuration and ignores the tool-use request.
Performance and Cost Considerations
Tool use incurs additional API latency and token costs. The architecture requires at least two LLM calls: an initial prompt and a final summarization call after tool execution, plus potential intermediate rounds for complex queries. Enable this feature only when the precision of graph-aware answers justifies the increased computational overhead.
Code Examples
Command-Line Usage
Run the reasoner from the terminal with or without tool use:
# Standard reasoning without graph tools
python src/reasoner.py \
--dataset flowvqa \
--reasoner gpt-4o \
--textualizer Qwen2-VL-7B \
--input_type mermaid
# Enable tool use for structural queries
python src/reasoner.py \
--dataset flowvqa \
--reasoner gpt-4o \
--textualizer Qwen2-VL-7B \
--input_type mermaid \
--tool_use
Programmatic API Access
Integrate tool use directly in Python by passing the representation to ModelWrapper:
from models.model_loader import ModelWrapper
from prompts import load_reasoner_prompt
# Prepare inputs
question = "How many nodes are in the flowchart?"
representation = "graph TD; A-->B; A-->C;"
prompt = load_reasoner_prompt(question, representation)
# Enable tool use by providing representation
model = ModelWrapper("gpt-4o")
answer = model.generate_response(prompt, representation=representation)
print(answer)
Advanced Manual Invocation
For custom implementations, access the tool-calling logic directly:
from models.api_models import generate_api_response_tool_use, load_tools
from openai import OpenAI
# Configure messages
messages = [
{"role": "system", "content": "You analyze flowcharts."},
{"role": "user", "content": "Count the nodes."}
]
# Execute with tool support
response = generate_api_response_tool_use(
model_name="gpt-4o",
client=OpenAI(api_key="YOUR_KEY"),
messages=messages,
representation="graph TD; Start-->End;"
)
print(response)
Summary
- Tool use requires the
--tool_useflag insrc/reasoner.pyand a valid Mermaid representation to activate the graph-analysis pipeline. - The
ModelWrapperclass insrc/models/model_loader.pydetects tool-use eligibility by checking for the presence of a representation argument. - Only
gpt-4oandgpt-4o-minisupport the tool-calling implementation found insrc/models/api_models.py. - Tools execute via
exec()on dynamically generated Python code that interfaces with theflowchartobject fromsrc/models/mermaid_parser.py. - Enable tool use for complex structural queries where precision outweighs the additional API latency and token costs.
Frequently Asked Questions
Which LLM models support tool use in the Textual Reasoner?
Only gpt-4o and gpt-4o-mini currently implement the tool-use pathway. The generate_api_response_tool_use function in src/models/api_models.py specifically handles these models, while other configurations fall back to standard text generation.
How does the system decide when to execute tool calls?
The LLM receives function definitions from load_tools() in src/models/prompt_utils.py and decides independently whether to return text or tool calls. If the model returns tool_calls, the system executes the corresponding Python code via exec() and feeds results back for final answer generation.
What performance impact does tool use have?
Tool use increases latency and token consumption because it requires multiple API calls: an initial request, tool execution rounds, and a final summarization request. According to the implementation in src/models/api_models.py (lines 86-119), each tool batch requires a separate round-trip to the LLM.
Can I add custom graph analysis tools?
While the repository provides standard tools like get_number_of_nodes through load_tools() in src/models/prompt_utils.py, extending the system would require modifying the function specifications in that file and ensuring the flowchart object in src/models/mermaid_parser.py exposes corresponding methods.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →