9 Atomic Tools in GenericAgent: Complete Implementation Guide in ga.py
The 9 atomic tools in GenericAgent are code_run, ask_user, web_scan, web_execute_js, file_patch, file_write, file_read, update_working_checkpoint, and start_long_term_update, implemented as do_* methods in ga.py spanning lines 296 to 508.
The GenericAgent framework (lsdefine/GenericAgent) exposes a minimal, secure boundary between large language models and the host environment. All external interaction—whether executing code, manipulating files, browsing the web, or persisting memory—flows through nine immutable primitive operations defined in ga.py. These atomic tools ensure deterministic behavior and comprehensive auditability.
Overview of the 9 Atomic Tools in GenericAgent
The toolkit divides naturally into five functional domains: code execution, user interaction, browser automation, filesystem operations, and memory management. Each tool is stateless and atomic, meaning it either completes entirely or fails without side effects, simplifying error recovery and retry logic.
| Tool | Domain | Core Purpose |
|---|---|---|
code_run |
Execution | Run Python, Bash, or PowerShell with timeout and output streaming |
ask_user |
Interaction | Halt execution for human input with optional candidate selection |
web_scan |
Browser | Capture simplified HTML or tab lists from the controlled browser |
web_execute_js |
Browser | Inject and return arbitrary JavaScript execution results |
file_patch |
Filesystem | Atomic find-and-replace within a file using unique content blocks |
file_write |
Filesystem | Create, overwrite, append, or prepend entire files |
file_read |
Filesystem | Read with line ranges, keyword search, and memory-access logging |
update_working_checkpoint |
Memory | Write temporary working memory to guide immediate next steps |
start_long_term_update |
Memory | Trigger durable knowledge extraction and SOP updates |
Implementation Details in ga.py
Each atomic tool maps to a do_* method inside ga.py. The GenericAgentHandler class implements these methods between lines 296 and 508, with the dispatcher automatically stripping the do_ prefix when routing LLM requests.
code_run Implementation
The do_code_run method (starting at line 296) handles safe execution of arbitrary code snippets. It supports Python, PowerShell, and Bash, injecting timeout controls and streaming output capture. The implementation writes temporary files when necessary and references assets/code_run_header.py for Python template initialization.
ask_user Implementation
Defined at line 320, do_ask_user constructs an interrupt payload that pauses the agent loop until the human operator responds. It supports optional candidate lists, forcing the user to select from predefined options rather than free-form input.
web_scan Implementation
The do_web_scan method (line 327) interfaces with the browser controller to retrieve simplified HTML snapshots. It can enumerate open tabs, switch to specific tab IDs, and return text-only content to reduce token consumption. The heavy lifting is delegated to simphtml.py.
web_execute_js Implementation
At line 342, do_web_execute_js injects arbitrary JavaScript into the controlled browser context. It captures return values, execution errors, and DOM mutations. The method optionally persists the JS return value to a specified file path.
file_patch Implementation
The do_file_patch method (line 370) performs atomic in-place file modifications. It searches for a unique old_content block and replaces it with new_content, including safety checks for missing or ambiguous matches to prevent corrupted states.
file_write Implementation
Starting at line 384, do_file_write handles whole-file operations: overwrite, append, or prepend. It accepts content via <file_content> tags or fenced code blocks and implements safe buffering for large writes.
file_read Implementation
The do_file_read method (line 418) provides flexible file access with optional line ranges (start, count), keyword search, and line-number display. It also logs memory-access events for files under the memory directory to track working set changes.
update_working_checkpoint Implementation
At line 447, do_update_working_checkpoint writes to the agent’s temporary working memory. It stores key_info and related_sop references that influence the immediate next steps without persisting to long-term storage.
start_long_term_update Implementation
The final tool, do_start_long_term_update (line 508), initiates a durable memory distillation process. When the agent identifies valuable task knowledge, this method prepares a prompt for the LLM to extract permanent facts and updates the memory SOP files.
How to Invoke the Tools: JSON Payload Examples
The GenericAgentHandler routes tool calls by stripping the do_ prefix from method names. Below are the canonical JSON payloads the LLM emits to invoke each atomic tool.
Execute Python with timeout:
{
"tool_name": "code_run",
"args": {
"type": "python",
"code": "print('Hello from GenericAgent')",
"timeout": 30
}
}
Request human input with candidates:
{
"tool_name": "ask_user",
"args": {
"question": "Which repository should I clone?",
"candidates": ["repo1", "repo2", "repo3"]
}
}
Capture browser state:
{
"tool_name": "web_scan",
"args": {
"tabs_only": false,
"switch_tab_id": "tab-3",
"text_only": true
}
}
Inject JavaScript:
{
"tool_name": "web_execute_js",
"args": {
"script": "document.title",
"save_to_file": "page_title.txt"
}
}
Atomic file modification:
{
"tool_name": "file_patch",
"args": {
"path": "config.yaml",
"old_content": "debug: false",
"new_content": "debug: true"
}
}
Write or append files:
{
"tool_name": "file_write",
"args": {
"path": "notes.txt",
"mode": "append",
"content": "<file_content>\nNew observation at $(date)\n</file_content>"
}
}
Read with search and pagination:
{
"tool_name": "file_read",
"args": {
"path": "README.md",
"start": 1,
"count": 100,
"keyword": "Installation",
"show_linenos": true
}
}
Update working memory:
{
"tool_name": "update_working_checkpoint",
"args": {
"key_info": "User prefers Chinese language output",
"related_sop": "sop/translation.md"
}
}
Trigger long-term memory consolidation:
{
"tool_name": "start_long_term_update",
"args": {}
}
Key Supporting Files
While ga.py contains the tool implementations, several adjacent modules provide essential infrastructure:
-
agent_loop.py: Defines theBaseHandlerclass and the orchestration loop that drives LLM-to-tool interaction, dispatching requests to thedo_*methods inga.py. -
simphtml.py: Provides HTML simplification and JavaScript execution capabilities for theweb_scanandweb_execute_jstools, bridging the agent to the browser controller. -
assets/code_run_header.py: Template header injected into temporary Python scripts executed bydo_code_run, setting up safe execution contexts. -
memory/global_mem_insight.txt: Static global memory context appended to prompts viaget_global_memory, providing persistent SOP guidance across agent sessions.
Summary
The 9 atomic tools in GenericAgent provide a complete, sandboxed interface between LLM reasoning and environmental interaction:
code_run: Safe execution of Python, Bash, or PowerShell with timeout controls.ask_user: Interactive human-in-the-loop interrupt with optional candidate selection.web_scan: Browser state inspection and tab management via simplified HTML capture.web_execute_js: Arbitrary JavaScript injection with return value capture.file_patch: Atomic find-and-replace for safe in-place file modification.file_write: Whole-file write, append, or prepend operations.file_read: Flexible file reading with search, pagination, and line-number display.update_working_checkpoint: Temporary working memory updates for immediate context.start_long_term_update: Durable memory distillation and SOP persistence.
Frequently Asked Questions
What distinguishes the 9 atomic tools in GenericAgent from other agent frameworks?
The GenericAgent architecture treats these nine operations as indivisible primitives, ensuring that every environmental interaction flows through a single, auditable channel in ga.py. Unlike frameworks that expose broad API surfaces, GenericAgent restricts the LLM to these specific do_* methods (lines 296–508), enabling precise sandboxing and deterministic replay of agent actions.
How does the code_run tool handle security and timeout constraints?
According to the ga.py implementation starting at line 296, do_code_run enforces security through temporary file isolation and explicit timeout parameters. The method injects a safe header from assets/code_run_header.py into Python executions, streams output back to the agent loop, and terminates processes that exceed the specified timeout, preventing runaway computations.
What is the difference between update_working_checkpoint and start_long_term_update?
update_working_checkpoint (line 447) writes temporary context to the agent’s working memory for immediate tactical guidance, such as storing user preferences for the current session. In contrast, start_long_term_update (line 508) initiates a knowledge distillation process that extracts durable facts from the completed task and persists them to the memory SOP, affecting future agent behavior across sessions.
Which files support the browser automation tools web_scan and web_execute_js?
While the tool definitions reside in ga.py (lines 327 and 342), the actual browser interaction logic relies on simphtml.py for HTML simplification and JavaScript execution context management. Additionally, agent_loop.py provides the orchestration layer that handles the asynchronous callback structure when these tools interrupt the agent flow to communicate with the browser controller.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →