How Running Agents Discover and Communicate in Prime Agent: The Daemon-Supervisor Architecture Explained

Prime Agent uses a centralized daemon-supervisor process combined with Unix-domain sockets and a registry directory to let agents discover peers and exchange messages without hard-coded network addresses.

Prime Agent implements a plug-and-play agent network where primary instances, sub-agents, and worker processes coordinate through a single supervisor. This architecture eliminates the need for static IP addresses or persistent peer-to-peer connections, enabling dynamic agent lifecycle management. Understanding how agent discovery and communication works is essential for building reliable multi-agent workflows with PrimeIntellect-ai/prime-agent.

The Daemon-Supervisor: Central Registry and Message Broker

Prime Agent's coordination layer centers on a daemon-supervisor process that fulfills two critical roles: maintaining an authoritative registry of all active agents and routing messages between them. The supervisor exposes its services through environment variables and a Unix-domain socket, creating a local-only communication fabric.

The supervisor watches a designated registry directory for agent registration records. Each agent writes a session-specific entry on startup, allowing the supervisor to build and maintain an in-memory map of the entire agent population. This directory path is controlled by the PRIME_AGENT_INTERNAL_DAEMON_SUPERVISOR_REGISTRY_DIR environment variable, keeping discovery configuration external to agent code.

Agent Discovery via the Registry Directory

When any agent launches—whether the primary Prime Agent instance, a sub-agent created through rlm.create_subagent(), or a worker process—it performs a registration handshake with the supervisor.

The discovery flow works as follows:

  • The agent writes a registration record to the supervisor's registry directory
  • The supervisor detects the new entry through filesystem watching and updates its internal agent map
  • Agents query this map through the daemon-client API to enumerate peers by their unique session IDs

This file-based approach provides durable registration without requiring network broadcasts or service discovery protocols. Agents that crash or exit cleanly are automatically removed from the active set when their registration records expire or are cleaned up.

The Daemon Protocol: Structured Communication Over Unix Sockets

All runtime communication between agents flows through a Unix-domain socket whose path is exposed via PRIME_AGENT_INTERNAL_DAEMON_SUPERVISOR_SOCKET. The socket implements the daemon protocol (version 7, schema revision ≥ 16), defined in packages/coding-agent/src/modes/daemon/daemon-protocol.ts.

Messages are JSON objects with standardized shapes covering:

  • heartbeat — Periodic keepalives to detect agent health
  • create_subagent — Requests to spawn new worker processes
  • list_subagents — Queries for the current agent population
  • session_event — Broadcast notifications across agent groups
  • run_tool and command dispatch — Targeted execution requests

The protocol separates concerns by message type: control plane operations (lifecycle management) share the same transport as data plane operations (tool execution and result streaming).

Creating and Communicating with Sub-Agents

Sub-agent interaction demonstrates the full discovery and communication stack in action. A parent agent initiates creation through the RLM runtime interface:

// In a running Prime Agent session (RLM runtime)
await rlm.create_subagent({
  name: "code-review",
  model: "gpt-4o-mini",
  prompt: "You are a code reviewer."
});

The request travels through daemon-client.ts to the supervisor, which:

  1. Spawns a new worker process via agent-runtime-host.ts
  2. Assigns a unique session ID and writes the registration record
  3. Returns a handle to the parent

The parent then addresses messages to the sub-agent using its session ID in the target field:

const targetId = "sub-agent-1234";
await rlm.send_message({
  target: targetId,
  command: "run_tool",
  payload: { tool: "eslint", args: ["src/**/*.ts"] }
});

Message Routing Architecture

Prime Agent implements a three-tier routing pattern:

Client → Supervisor — Agents use the daemon client (daemon-client.ts) to submit commands. The client handles serialization, socket connection management, and response correlation.

Supervisor → Worker — The supervisor inspects message targets and forwards to specific workers, or broadcasts to all workers for session-wide events. Routing decisions use the in-memory agent map built from registry directory state.

Worker → Supervisor — Workers emit status updates and results through the same socket. Events like subagent_status and session_event are relayed to interested peers based on subscription patterns.

This symmetric design means any agent can initiate communication with any other discovered agent without direct socket connections between peers.

Capability Negotiation and Version Compatibility

The supervisor advertises supported features through DAEMON_DEFAULT_SERVER_CAPABILITIES in daemon-protocol.ts. Clients validate capabilities before issuing optional commands, enabling graceful degradation when connecting to supervisors with different feature sets.

Current capabilities include:

  • delete_rlm_subagent — Explicit sub-agent termination
  • model_catalog — Runtime model availability queries

This contract ensures that agent code written against newer protocol versions fails predictably rather than silently when running against older supervisors.

Heartbeat and Health Monitoring

The protocol includes explicit heartbeat messages for failure detection. The daemon client automatically emits periodic heartbeats:

// daemon-client.ts – sends periodic heartbeats
setInterval(() => {
  client.send({ type: "heartbeat", sessionId: mySessionId });
}, 5_000);

Supervisors track heartbeat timestamps per agent and can declare agents failed when heartbeats lapse, triggering cleanup of registry entries and notification to interested peers.

Environment Configuration Reference

Prime Agent's discovery and communication are entirely configured through environment variables:

Variable Purpose
PRIME_AGENT_INTERNAL_DAEMON_SUPERVISOR_REGISTRY_DIR Filesystem path for agent registration records
PRIME_AGENT_INTERNAL_DAEMON_SUPERVISOR_SOCKET Unix-domain socket path for protocol messages
PRIME_AGENT_INTERNAL_DAEMON_WORKER_PROTOCOL_SOCKET Worker-specific socket override in daemon-worker-protocol.ts

These variables are set by the launcher (daemon-update-restart.ts) and inherited by all child processes, ensuring consistent addressing across the agent hierarchy.

Summary

  • Discovery happens through a watched registry directory where agents write session records; the supervisor maintains the authoritative agent map
  • Communication uses Unix-domain sockets with a versioned JSON protocol supporting heartbeats, commands, and events
  • Sub-agents are created through rlm.create_subagent() and addressed by session ID via the same socket infrastructure
  • Capability gating in DAEMON_DEFAULT_SERVER_CAPABILITIES ensures protocol compatibility
  • All configuration flows through environment variables set by the launcher, keeping network addresses out of agent code

Frequently Asked Questions

How does Prime Agent handle agent failures?

The supervisor detects failed agents through missed heartbeats. When a heartbeat deadline passes, the supervisor removes the agent's registry entry and emits session_event notifications to subscribers. This allows parent agents to implement retry logic or sub-agent replacement without polling.

Can agents communicate across machine boundaries?

No. The current architecture uses Unix-domain sockets and local filesystem watching, restricting agent communication to a single host. Multi-host coordination would require a federation layer outside the current daemon-supervisor.ts implementation.

What limits the number of concurrent sub-agents?

Practical limits come from filesystem watchers on the registry directory (kernel-dependent, typically thousands) and socket file descriptor availability. The protocol itself imposes no hard caps; the supervisor's in-memory agent map scales with available RAM.

How do I debug discovery failures?

Check that PRIME_AGENT_INTERNAL_DAEMON_SUPERVISOR_REGISTRY_DIR and PRIME_AGENT_INTERNAL_DAEMON_SUPERVISOR_SOCKET are set identically across parent and child processes. Verify the supervisor process is running and listening with lsof -U | grep prime. Examine daemon-supervisor.ts logs for registration record parse errors or permission failures on the registry directory.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →