# How Tool Use Works in the Textual Reasoner: Implementation and Best Practices

> Unlock Textual Reasoner's tool use to let LLMs execute graph analysis on Mermaid flowcharts. Learn how it works and when to enable this powerful feature for accurate structural queries.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: best-practices
- Published: 2026-03-05

---

**Tool use in the Textual Reasoner enables supported LLMs to execute graph-analysis functions on Mermaid flowcharts by passing the `--tool_use` flag, triggering a multi-step pipeline where the model can invoke Python functions like `get_number_of_nodes` to answer structural queries accurately.**

The Textual Reasoner in the **junyiye/textflow** repository provides an optional tool use capability that allows large language models (LLMs) to perform precise graph analysis on flowchart representations. When enabled, this feature transforms natural language questions about flowchart structure into executable function calls, bridging the gap between textual reasoning and computational graph operations.

## How Tool Use Works in the Textual Reasoner

### Enabling the Tool Use Flag

In [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), the system exposes a boolean command-line flag `--tool_use` (lines 42-45) that initiates the tool-calling pathway. When this flag is present, the argument parser stores `True` in `args.tool_use`, signaling that the incoming Mermaid representation should be treated as a traversable graph object rather than plain text.

### Flag Propagation to ModelWrapper

The flag value propagates through the inference pipeline in [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) (lines 98-105). If `tool_use` is enabled, the `ModelWrapper.generate_response` method receives both the text prompt and the `representation` argument (containing the Mermaid code). Without the flag, only the prompt is forwarded, bypassing the tool-use logic entirely.

### Detection Logic in ModelWrapper

Inside [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py), the `generate_response` method (lines 28-34) automatically detects tool-use eligibility by checking whether a `representation` was supplied. The presence of a representation implicitly toggles `tool_use = True`, ensuring that graph data is available when the LLM attempts function calls, even if the explicit flag was omitted.

### The LLM Tool-Calling Implementation

For API-based models including `gpt-4o` and `gpt-4o-mini`, the system invokes `generate_api_response_tool_use` from [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py) (lines 59-66). This function implements a two-phase protocol: it first sends the prompt without tools, and only if the LLM returns `null` content does it fall back to the tool-calling pathway.

### Tool Definitions and Execution Loop

The available tools are defined in [`src/models/prompt_utils.py`](https://github.com/junyiye/textflow/blob/main/src/models/prompt_utils.py) (lines 13-30) via `load_tools()`, which returns OpenAI function-calling schema specifications for operations like `get_number_of_nodes` and `get_direct_successors`. When the LLM returns `tool_calls`, [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py) (lines 86-119) executes `load_tools_code()` to generate Python code that forwards calls to the `flowchart` object. Each call runs via `exec()`, with results captured and appended as tool messages. A second LLM request then processes these results alongside the original prompt to generate the final answer.

## When to Enable Tool Use in the Textual Reasoner

Enable the `--tool_use` flag when your workflow meets specific criteria regarding query complexity, model compatibility, and performance tolerance.

**Complex Graph Queries**

Activate tool use when downstream questions require structural computation that natural language processing cannot reliably perform, such as counting nodes ("How many nodes are there?"), traversing paths ("What is the shortest path between X and Y?"), or analyzing connectivity. The explicit function calls eliminate ambiguity in numerical or topological answers.

**Supported LLM Models**

Currently, only `gpt-4o` and `gpt-4o-mini` implement the tool-use pathway in [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py). Enabling the flag with unsupported models triggers a silent fallback to standard text generation, as the `ModelWrapper` detects the unsupported configuration and ignores the tool-use request.

**Performance and Cost Considerations**

Tool use incurs additional API latency and token costs. The architecture requires at least two LLM calls: an initial prompt and a final summarization call after tool execution, plus potential intermediate rounds for complex queries. Enable this feature only when the precision of graph-aware answers justifies the increased computational overhead.

## Code Examples

### Command-Line Usage

Run the reasoner from the terminal with or without tool use:

```bash

# Standard reasoning without graph tools

python src/reasoner.py \
  --dataset flowvqa \
  --reasoner gpt-4o \
  --textualizer Qwen2-VL-7B \
  --input_type mermaid

# Enable tool use for structural queries

python src/reasoner.py \
  --dataset flowvqa \
  --reasoner gpt-4o \
  --textualizer Qwen2-VL-7B \
  --input_type mermaid \
  --tool_use

```

### Programmatic API Access

Integrate tool use directly in Python by passing the representation to `ModelWrapper`:

```python
from models.model_loader import ModelWrapper
from prompts import load_reasoner_prompt

# Prepare inputs

question = "How many nodes are in the flowchart?"
representation = "graph TD; A-->B; A-->C;"
prompt = load_reasoner_prompt(question, representation)

# Enable tool use by providing representation

model = ModelWrapper("gpt-4o")
answer = model.generate_response(prompt, representation=representation)
print(answer)

```

### Advanced Manual Invocation

For custom implementations, access the tool-calling logic directly:

```python
from models.api_models import generate_api_response_tool_use, load_tools
from openai import OpenAI

# Configure messages

messages = [
    {"role": "system", "content": "You analyze flowcharts."},
    {"role": "user", "content": "Count the nodes."}
]

# Execute with tool support

response = generate_api_response_tool_use(
    model_name="gpt-4o",
    client=OpenAI(api_key="YOUR_KEY"),
    messages=messages,
    representation="graph TD; Start-->End;"
)
print(response)

```

## Summary

- Tool use requires the `--tool_use` flag in [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) and a valid Mermaid representation to activate the graph-analysis pipeline.
- The `ModelWrapper` class in [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py) detects tool-use eligibility by checking for the presence of a representation argument.
- Only `gpt-4o` and `gpt-4o-mini` support the tool-calling implementation found in [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py).
- Tools execute via `exec()` on dynamically generated Python code that interfaces with the `flowchart` object from [`src/models/mermaid_parser.py`](https://github.com/junyiye/textflow/blob/main/src/models/mermaid_parser.py).
- Enable tool use for complex structural queries where precision outweighs the additional API latency and token costs.

## Frequently Asked Questions

### Which LLM models support tool use in the Textual Reasoner?

Only `gpt-4o` and `gpt-4o-mini` currently implement the tool-use pathway. The `generate_api_response_tool_use` function in [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py) specifically handles these models, while other configurations fall back to standard text generation.

### How does the system decide when to execute tool calls?

The LLM receives function definitions from `load_tools()` in [`src/models/prompt_utils.py`](https://github.com/junyiye/textflow/blob/main/src/models/prompt_utils.py) and decides independently whether to return text or tool calls. If the model returns `tool_calls`, the system executes the corresponding Python code via `exec()` and feeds results back for final answer generation.

### What performance impact does tool use have?

Tool use increases latency and token consumption because it requires multiple API calls: an initial request, tool execution rounds, and a final summarization request. According to the implementation in [`src/models/api_models.py`](https://github.com/junyiye/textflow/blob/main/src/models/api_models.py) (lines 86-119), each tool batch requires a separate round-trip to the LLM.

### Can I add custom graph analysis tools?

While the repository provides standard tools like `get_number_of_nodes` through `load_tools()` in [`src/models/prompt_utils.py`](https://github.com/junyiye/textflow/blob/main/src/models/prompt_utils.py), extending the system would require modifying the function specifications in that file and ensuring the `flowchart` object in [`src/models/mermaid_parser.py`](https://github.com/junyiye/textflow/blob/main/src/models/mermaid_parser.py) exposes corresponding methods.