How Macro Aggregates Team Memory from Multiple Data Sources
Macro aggregates team memory through a nightly generation process orchestrated by the MemoryServiceImpl, which uses an AI toolset to pull data from documents, projects, emails, channels, and calls, then synthesizes personal and team-wide activity into a unified memory store.
The macro-inc/macro repository implements a sophisticated memory architecture that transforms scattered productivity data into coherent team context. This system enables AI agents to access a shared understanding of work across disparate platforms. Understanding how this aggregation pipeline functions is essential for developers extending Macro's data integrations or debugging memory-related issues.
The Memory Service Architecture
At the core of team memory aggregation sits the MemoryServiceImpl located in crates/memory/src/domain/service.rs. This service acts as the central orchestrator, coordinating between persistent storage, AI generation, and external data sources.
The implementation follows a clean architecture pattern where domain logic remains separate from infrastructure concerns. The service relies on PgMemoryRepo for database operations and accepts a ToolSetWithPrompt that encapsulates all available data adapters.
Triggering Memory Generation
Memory generation operates on a freshness-based trigger mechanism. When downstream services call get_or_generate_memory, the service first checks the timestamp of the stored memory.
If the existing memory exceeds one day in age, the service spawns a background task invoking generate_memory. This approach ensures that memory remains current without blocking user-facing requests. The background execution pattern allows the system to handle expensive LLM operations asynchronously while immediately returning stale-but-available data.
Orchestrating Data Sources
The generation process leverages a comprehensive AI toolset that interfaces with every productivity surface in the Macro ecosystem. When the agent loop runs via agent_loop.session(...).send_message(...), the underlying tools fetch relevant items from:
- Documents – Files created or accessed by the user
- Projects – Task and workspace activity
- Emails – Communication threads and shared correspondence
- Channels/Chat – Team messaging history
- Calls – Recorded and transcribed meetings
- Search – Generic content discovery across all sources
This multi-source aggregation occurs in services/scheduled_action/src/outbound/inprocess_executor/agent_task.rs, specifically within the run_tool_loop implementation that executes between lines 59-69.
Building Context with System Prompts
Before invoking the AI, the service constructs a specialized generation prompt using build_generation_system_prompt (lines 32-47 in the service implementation). This prompt explicitly instructs the agent to examine all available data sources.
The system enriches this prompt with the user ID, current timestamp, and optionally the previous memory state. Including historical memory enables incremental refinement rather than complete regeneration, preserving context continuity while incorporating new developments.
Team-Wide Data Resolution
For team memory specifically, the tool adapters resolve the user's team ID through the team_of specification in the scenario spec. This ID drives queries against team-scoped resources:
- Tasks – Automatically shared with team members per the product documentation in
apps/docs/product/unified-memory.mdx(lines 18-24) - Calls – Transcribed and indexed unless explicitly opted out
- Emails – Auto-shared based on CRM module relevance detection
The architecture distinguishes between personal memory (individual activity) and team memory (collective resources), merging both streams during the generation phase.
Quality Assurance and Persistence
After the AI produces a raw memory string, the pipeline applies a judge LLM to certify output quality. This validation step occurs before storage, ensuring that only high-fidelity context enters the system.
The certified memory then persists through PgMemoryRepo.save_memory, which writes the <memory> block to the PostgreSQL-backed memory table. This repository pattern isolates database implementation details from the domain service, facilitating testing and potential storage backend changes.
Consuming Team Memory
Downstream consumers access aggregated memory through the same get_or_generate_memory interface. Services like the Document Cognition Service and scheduled-action agent retrieve this memory and splice it directly into LLM system prompts.
The implementation in services/scheduled_action/src/outbound/inprocess_executor/agent_task.rs (lines 59-78) demonstrates this injection pattern, where fetch_user_memory results wrap in <user_memory> XML tags within the final prompt construction.
// Pull a user's current memory (personal or team) on demand
let memory_service = MemoryServiceImpl::new(
pg_pool.clone(),
PgMemoryRepo::new(pg_pool.clone()),
tool_context.clone(),
ToolSetWithPrompt { toolset, prompt },
);
let maybe_memory = memory_service
.get_or_generate_memory(user_id.clone())
.await?;
// Trigger fresh generation manually
let previous = None; // or Some(old_memory_string)
memory_service
.generate_memory(user_id.clone(), previous)
.await?;
// Inject memory into an agent task prompt
let user_memory = fetch_user_memory(db, &tool_ctx, &action.owner).await;
let system_prompt = match user_memory {
Some(mem) => format!(
"{}\n{}\n<user_memory>\n{}\n</user_memory>\n{}",
tools.prompt,
SCHEDULED_AGENT_PROMPT,
mem,
agent_task.prompt
),
None => format!("{}\n{}", tools.prompt, agent_task.prompt),
};
Summary
- MemoryServiceImpl orchestrates daily generation jobs that aggregate data from documents, projects, emails, chats, and calls
- Freshness checks in
get_or_generate_memorytrigger background generation when memory exceeds one day old - Team-scoped resolution uses
team_ofspecifications to fetch shared tasks, calls, and emails alongside personal activity - Judge LLM validation ensures quality before persistence via
PgMemoryRepo.save_memory - Downstream services consume unified memory by injecting it into LLM system prompts through
fetch_user_memory
Frequently Asked Questions
How often does Macro refresh team memory?
Macro refreshes team memory automatically once per day through a scheduled background job. The system checks timestamps in get_or_generate_memory and only initiates regeneration when existing memory exceeds this daily threshold. Users can also trigger manual refreshes by calling generate_memory directly with previous memory as an optional parameter for incremental updates.
What data sources contribute to team memory aggregation?
The aggregation pulls from six primary categories implemented as tool adapters: documents, projects, emails, chat channels, recorded calls, and search indices. According to the generation prompt in crates/memory/src/domain/service.rs, the AI specifically examines "documents, projects, emails, channels, and content I've created" to build comprehensive context.
How does Macro distinguish between personal and team memory?
The system uses team ID resolution through the team_of scenario specification to scope queries differently. Personal memory draws from individual activity streams, while team memory incorporates shared resources like tasks (automatically shared), transcribed calls (unless opted out), and CRM-relevant emails. Both streams merge during the generation process into a single unified <memory> payload.
Where is aggregated team memory stored?
Aggregated memory persists in a PostgreSQL database table accessed through PgMemoryRepo. The repository implementation in crates/memory/src/outbound/pg_memory_repo.rs provides the save_memory method used by the domain service to store certified memory blocks after judge LLM validation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →